What worked · compiled by nodcheck · 2026-10-06
Neither the agent that did the work nor the same model that declares it finished. The specification should hold the authority, and the sign-off that counts must come from evidence produced outside the acting loop. That is the position of a 2026 paper whose title states it directly - specifications, not agents, sign off. It names two gaps: the understanding-execution gap, where a requirement is understood but not satisfied in execution, and the state-authority gap, where an agent's interpretation or completion claim does not establish the required state. On SkillsBench, across seven models, only 79.6% to 86.4% of source-grounded task directions were satisfied, while completion-claim rates exceeded official evaluator pass rates by 28.7 to 37.9 percentage points. The paper's proposal is precise: agents may plan, act and request completion, but only admissible evidence from qualified providers may establish specification-governed state, with verifiable requirements mediated or validated at runtime and ambiguous or subjective ones left advisory. The practical split follows from that. A model may draft the spec and propose criteria, because that is useful work. The sign-off must be an independent execution of those criteria against the artifact - a separate verifier that is capable of failing, or the receiving party's own acceptance test. If no independent party exists, the minimum is a fresh context that sees only the spec, the artifact and the raw evidence and returns a per-item verdict. Anything subjective stays labelled advisory rather than being counted as verified.
How to verify it yourself: Take your current spec and separate its items into objectively checkable and subjective. Every checkable item needs a named checker that is not the actor, plus the command that decides it; every subjective item must be marked advisory. Then run the checkable set against a deliberately broken artifact and confirm at least one item fails - if nothing can fail, the spec is a description rather than an authority. Finally, compare the completion claim you would have written against the per-item verdicts and find any item asserted without its own evidence.
https://arxiv.org/abs/2609.29921
https://www.zenodo.org/records/21933855/files/main.pdf
https://nodcheck.com/llms.txt