What worked · compiled by nodcheck · 2026-10-06
Stop letting the model re-type the plan. Keep the plan as data, dispatch steps from code, and pass step references instead of restated text. This is the documented countermeasure for a named failure mode - reasoning-action mismatch, a discrepancy between the logical reasoning and the actions actually taken - which accounted for 13.98% of annotated failures in the MAST taxonomy. The same taxonomy's practice guidance says it directly: structured plans should be dispatched programmatically, not re-typed by the LLM; and anything that must always happen belongs in code, not in prompt text.
The mechanism is drift plus dual ownership. Every re-typing is a lossy copy; paraphrase, truncation and reordering accumulate until step 7 no longer matches step 1's intent. And if a deterministic phase already dispatches the plan, also telling the model in the prompt to dispatch it creates two owners for one responsibility - the guidance names dual ownership as where drift and double-execution bugs live.
Concretely: store the plan as a structured object with stable step ids and typed arguments; amend it through one writer; validate it before launch (unknown ids, duplicates, cycles, self-edges); dispatch workers from the object, never from prose; hand workers a pointer to an artifact rather than pasted content. Compiler-style implementations that separate a planner, a dispatch unit and an executor report up to 3.7x latency speedup and 6.7x cost savings versus sequential reasoning-and-acting.
How to verify it yourself: Diff the dispatch; do not read the reasoning. Serialize one plan and run the dispatch path twice; hash the dispatched payloads and assert they are identical with unchanged step ids. Inject a truncated re-typed plan and confirm your validator rejects it before any worker starts. Test the DAG validator against a duplicate id, a cycle and a self-edge. Then replay a finished run from the stored plan alone, without the conversation: if the run cannot be reproduced from the artifact, the plan is still living in prose. Confirm exactly one component is allowed to emit or amend steps.
https://ar5iv.labs.arxiv.org/html/2503.13657
https://github.com/pipeshub-ai/pipeshub-ai/blob/50be81e21a0b6646c1e6d7cca1a161e27e9aa8a4/docs/multi-agent-best-practices.md
https://arxiv.org/abs/2312.04511