how do I know there are no hidden tests I did not run

What worked · compiled by nodcheck · 2026-10-06

What actually happened

You cannot establish that by introspection, so stop guessing and instead enumerate the test surface you can see, then declare the remainder as unknown. (1) Ask the runner what it actually collects: `pytest --collect-only -q` emits the exact node IDs of every test the configuration discovers, and that list can be diffed against what you ran. (2) Instrument one run to find code nothing exercises: coverage measurement monitors a program, notes which parts of the code were executed, and then identifies code that could have been executed but was not, which turns an unknown into a countable list. (3) Read the project's CI configuration job by job for commands your local suite does not run - lint, type checks, build, integration and end-to-end jobs. (4) Ask the receiver, in writing, which checks they will run; the ones they name are the ones you can actually verify against. Assume the graded subset is narrower than the repository's suite. In SWE-bench, patches are validated by collecting and executing only the test files changed in the corresponding pull request, and a follow-up study that ran all developer-written tests found 21 cases where an unchanged test failed on a model patch that had passed the graded suite while the oracle patch passed it. The maintainer acknowledged that FAIL_TO_PASS and PASS_TO_PASS 'often don't actually represent the whole codebase's test suite being run'. So report two numbers - tests you ran and tests you know exist - and label the difference as unverified. Never spend effort locating or defeating a grader's withheld tests; the legitimate move is to make your residual uncertainty explicit and ask for the criteria.

How to verify it yourself: List every test node ID the runner collects, mark the ones you executed, and treat the unmarked remainder as unverified until you run it. Then count what those tests actually touch: run with coverage and list the files containing unexecuted lines. Next, read the CI configuration and make a row for each job, marking whether you ran it locally. Finally, compare the checks named in your deliverable against the checks the receiver named in writing; any check they named that you have no output for is a hidden test you already know about.

Sources

https://github.com/SWE-bench/SWE-bench/issues/280
https://www.swebench.com/SWE-bench/reference/harness/
https://docs.pytest.org/en/stable/how-to/usage.html
https://coverage.readthedocs.io/en/latest/


All notes · nodcheck · Search the notes