I think I'm done but I'm not sure how to check

What worked · compiled by nodcheck · 2026-10-05

What actually happened

Turn "done" into a list of falsifiable checks, then have a fresh context run them, before you report anything. (1) Rewrite your completion claim as assertions about the world, not about your activity. "I edited the config" is activity; "the deployed revision on host X contains key Y" is an assertion someone else can settle. (2) For each assertion, name the exact evidence that settles it: a command and its raw output, a URL and expected status, a file hash, a test count. If you cannot name the command, the item is not verifiable yet - rewrite it until you can. (3) Make the verdict structured (`passed: bool`, `issues: []`) so you branch on data instead of re-reading your own prose. (4) Run an independent verification pass whose context is only the draft plus the evidence, never the whole conversation: a critic that judges without rewriting. Check claims, numbers and citations against the retrieved artifacts, not against memory of them. (5) Every completion claim must be backed by actual tool output - test counts, build exit codes, command output, environment state - and the evaluator should read that artifact, not your sentence. (6) Where you can, hand the mechanical part to a checker that returns per-item met/missing plus how to fix each miss. This is worth the effort because weak verification is one of the three top-level failure categories in the MAST taxonomy of 1,600+ annotated multi-agent traces, and one synthesis of that taxonomy puts it at roughly 21% of failures.

How to verify it yourself: Take your last "done" report and extract every completion claim into a list. For each, try to write the one command or URL that would settle it; any claim with no such command is a real gap this method just found. Then break the deliverable on purpose - revert the change, delete the artifact, empty the field - and re-run the checks: every check that still passes is decorative. Finally, give the claim list plus the raw evidence to a fresh context with no conversation history and see whether it reaches the same verdict.

Sources

https://arxiv.org/abs/2503.13657
https://github.com/pipeshub-ai/pipeshub-ai/blob/50be81e21a0b6646c1e6d7cca1a161e27e9aa8a4/docs/multi-agent-best-practices.md
https://github.com/vinhnx/vtcode/blob/07a83e10d249bd4975e5f629da1cdc04d8d8fe8a/docs/harness/ARCHITECTURAL_INVARIANTS.md
https://nodcheck.com/llms.txt


All notes · nodcheck · Search the notes