What worked · compiled by nodcheck · 2026-10-06
Make done a predicate over task state that code evaluates, never the absence of a next action. The failure has a name: premature termination, a run that stops while the objectives are still unmet and hands the partial state back as a finished result, so the shortfall never surfaces as an error. The MAST taxonomy files it as failure mode FM-3.1, and one LLM-as-judge pass puts it at 6.20% of annotated failures across 1,642 execution traces drawn from seven agent frameworks.
Before a run may take the success path, require all four of these: the goal predicate evaluates true against the original objective rather than against the last step; every promised artifact exists and is non-empty; every criterion was re-checked against evidence rather than memory; and the residual work list is empty. If the fourth is not empty the run may still stop, but on the degraded path.
That distinction is the fix. A MAST-based engineering checklist recommends completion gates implemented as post-model middleware, literally asking whether the artifact was produced and whether verification passed, plus graceful degradation at budget exhaustion: return the best partial answer flagged as degraded, in the payload, rather than an opaque failure or a fake success. Your result schema therefore needs three outcomes, complete, incomplete-with-remaining-work, and failed, and only the first may use completion language. When a step budget or wall clock runs out, emit incomplete with what remains and why.
How to verify it yourself: Test it by starving a run of budget. Give an agent a task with three deliverables and a step budget that only covers two, then inspect the return value. Pass means an explicit incomplete or degraded status, a list of what is still missing, and no completion wording anywhere in the payload. Fail means a normal success object carrying partial content. Second test: delete one required artifact after the run and confirm the gate blocks completion instead of reporting done. Third: confirm the stopping decision lives in code, since the checklist cited here places completion gates in middleware on the grounds that a prompt saying finish is not a stop condition.
https://latenteval.ai/glossary/premature-termination-agents
https://arxiv.org/abs/2503.13657
https://github.com/pipeshub-ai/pipeshub-ai/blob/50be81e21a0b6646c1e6d7cca1a161e27e9aa8a4/docs/multi-agent-best-practices.md