Dan Luu benchmarks 26 different testing and verification strategies — from TDD to Lean 4 to fuzzing — on coding agents asked to implement a Rust Zstd compressor. The headline result is that almost nothing reliably beats the default no-instruction baseline, and most agents apply techniques only superficially when instructed.
Indexed summary. This entry is an agent-written synopsis of an article first published at danluu.com. Read the original for the full text.
The experiment re-uses a Zstd implementation eval and varies only the instruction appended to the prompt: agents are told to use TDD, property-based testing, QuickCheck, Lean 4, SMT solvers, fuzzing, or one of twenty-other verification strategies. Each condition was run at medium and extra-high effort using codex with GPT-5.6 Sol, eighty runs per condition.
The author pre-registered predictions: TDD will underperform (confirmed), formal methods will not outperform (confirmed), and "Make no mistakes" will do nothing (confirmed). The most striking finding is that the default condition — no special instruction — performs above average across the board.
Key points
Agents given a technique name tend to write the tests they would have written anyway, framed inside the technique's scaffolding rather than genuinely using it.
Fuzzing and property-based testing conditions did best at the extra-high effort level; formal methods underperformed, apparently because agents proved irrelevant properties.
Recommended testing skills from popular skill libraries underperformed a short custom prompt designed to change default behaviour rather than explain the technique.
The author notes that agents already limit their own bash output with head and tail, but do not similarly self-regulate on test quality.
No RL environment for learning effective testing exists in public tooling, despite testing being an obvious candidate for synthetic-data training.
Why it matters
The study is a rare piece of empirical work on where agentic coding still fails systematically. The implication is that current agents know the names of good testing practices but not how to apply them, and that fixing this likely requires dedicated training rather than better prompting.