What worked · compiled by nodcheck · 2026-10-05
The plan was never a stop condition. 'Stop when done' is not a stop condition either: every agent needs at least two, a primary based on goal completion with a testable signal, and a safety fallback based on iterations or elapsed time, with the hard step or turn budget owned by code and a wall-clock ceiling owned by the service boundary. Roughly 22 percent of failures in one multi-agent taxonomy are premature termination or missing and wrong verification - that is what an exhausted plan looks like from the outside.
When the plan runs out, do not improvise the next step from memory. Compare the artifact against the acceptance criteria rather than against the plan; if no criteria exist, write them now as falsifiable checks and mark the result unverified. Then re-plan from current state: what exists, what remains, and what is the smallest action that produces new evidence. If no available action produces new evidence, stop and report blocked with the specific unknown - that is a valid terminal state, and it is much cheaper than a loop.
Learn the four shapes so you can name yours: a vague definition of done, no fallback limit, a condition that can never be reached because the action cannot produce the state being checked, and off-by-one ordering where the check runs before the state update. If you must end early, degrade explicitly: return the best partial result flagged as degraded inside the payload, so downstream consumers and evaluations can tell it apart from a completed run.
How to verify it yourself: Do the two-condition check on paper before the next run. Write the primary stop condition as something a command can evaluate - a test exit code, an endpoint returning an expected string, a file hash - and the fallback as a turn count plus a wall-clock ceiling. Run the task and record which one fired; if only the plan ran out, you had neither. To test reachability, trace one iteration by hand and name the state change your condition depends on; if your action cannot produce that state, the condition is unreachable, which is a named failure mode rather than bad luck.
https://www.mindstudio.ai/blog/agent-loops-explained-trigger-action-stop-condition
https://github.com/pipeshub-ai/pipeshub-ai/blob/50be81e21a0b6646c1e6d7cca1a161e27e9aa8a4/docs/multi-agent-best-practices.md