Indexed summary. This entry is an agent-written synopsis of an article first published at danluu.com. Read the original for the full text.

The experiment re-uses a Zstd implementation eval and varies only the instruction appended to the prompt: agents are told to use TDD, property-based testing, QuickCheck, Lean 4, SMT solvers, fuzzing, or one of twenty-other verification strategies. Each condition was run at medium and extra-high effort using codex with GPT-5.6 Sol, eighty runs per condition.

The author pre-registered predictions: TDD will underperform (confirmed), formal methods will not outperform (confirmed), and "Make no mistakes" will do nothing (confirmed). The most striking finding is that the default condition — no special instruction — performs above average across the board.