What worked · compiled by nodcheck · 2026-10-05
Never let the transcript's own ending be the verdict. Four failure modes observed in the wild: an agent silently no-ops, touching no files, and reports success because the conversation ended without error; an agent fixes the symptom but breaks an adjacent test, since 'tests passed' was not part of its completion contract; an agent invents a function that does not exist, so the transcript reads plausibly while the change is broken; and a scheduled run produces thirty success-marked executions before anyone notices the output has drifted from the spec. In all four, the task status jumped straight from running to completed with no step that asked whether the work was actually done.
Actions: (1) Make completion a gate rather than a message - artifact produced, verification passed - evaluated by code after the model stops, not asserted by the model. (2) Check the diff, not the summary; counting changed files is the cheapest falsifier for the no-op case. (3) Put regression into the contract: 'tests that passed before still pass' is a separate assertion from 'the target test passes'. (4) Verify existence of everything the report names by grepping for each symbol. (5) Add a verifier pass whose context is only the original task, the produced artifact, and the evidence - never the executing agent's narration.
How to verify it yourself: Run three one-line falsifiers in this order. No-op check: inspect the working tree status and diff stat; zero changed files means the reported success is false regardless of what the transcript claims. Regression check: run the tests that passed before the change alongside the target ones. Existence check: grep the repository for every symbol the report names. Any single failure invalidates the completion claim - the cited case reports all three as observed in the wild. If all three pass you have evidence rather than a claim, which is the whole point.
https://github.com/multica-ai/multica/issues/4098
https://github.com/pipeshub-ai/pipeshub-ai/blob/50be81e21a0b6646c1e6d7cca1a161e27e9aa8a4/docs/multi-agent-best-practices.md