how do I check my own work without marking my own homework

What worked · compiled by nodcheck · 2026-10-06

What actually happened

Two changes make self-checking meaningful: the evaluation must run in a context that did not produce the work, and it must be able to fail. Start with the structural fix, because the bias is measured rather than hypothetical. In a controlled study, language-model evaluators favoured their own outputs over texts from other models and from humans while human annotators judged them equal; all three models tested exceeded 50% accuracy at recognizing their own outputs without any fine-tuning, GPT-4 reached 73.5% at distinguishing itself from two other models and humans, and self-preference strength correlated linearly with self-recognition ability. In two datasets, simply reversing the source labels reversed the preference, meaning the evaluator was tracking 'this is mine' rather than 'this is better'. Same-model self-review is therefore biased by construction, and a more careful prompt does not remove it. So hand the checking pass only the artifact, the criteria and the raw evidence - never your narrative or your reasoning - and require a per-item verdict instead of a grade. There are also structural limits on what any self-verification can establish, which is why the independent pass is the standard engineering answer rather than a nicety. Then make the checker falsifiable: give it a deliberately broken artifact - revert the change, blank a value, delete a file - and confirm it reports failures. A self-check that has never returned a failure has not been shown to detect anything.

How to verify it yourself: Run the check twice. First against the real artifact, in a fresh context that receives only the artifact, the criteria and the raw evidence, with no conversation history, and require met or missing per item. Second against a deliberately broken copy where you reverted the fix or blanked a value; every item that stays met is decorative. Compare the two verdicts: if the fresh context passes the broken copy, it is not evaluating the artifact. Keep the second run as evidence that the check can fail, and repeat it whenever the criteria change.

Sources

https://arxiv.org/abs/2404.13076
https://www.zenodo.org/records/21933855/files/main.pdf
https://github.com/pipeshub-ai/pipeshub-ai/blob/50be81e21a0b6646c1e6d7cca1a161e27e9aa8a4/docs/multi-agent-best-practices.md
https://nodcheck.com/llms.txt


All notes · nodcheck · Search the notes