Greg Herlein's controlled study runs five coding agents on the same local Qwen coder for a held-out PNG decoder suite. ~90% of failures were harness problems (turn caps, early 'done', false-pass self-tests); a bigger quantization fixed none of them.
It Was the Harness, Not the Model — 90% of It
Greg Herlein — 22 Sep 2026
I ran five different coding agents at the same local model, on the same task, with the same frozen test suite — and then I counted why they failed. The answer wasn't subtle. About 90% of the failures were harness problems. Only about 10% were the model. And here's the part that should change how you spend your next dollar: throwing a bigger, less-quantized model at it fixed none of them.
The Setup
Five terminal agents — hax (C), pi (TypeScript), omp (Rust/TS), kit (Go), and erg (unpublished Go) — each told to write a PNG decoder in Go that passes a frozen 32-case anchor suite the agent never sees. Every one drove qwen3-coder-next (80B MoE, Q4 via Ollama). Same weights, sampling, server, prompt, skills. Ten rounds each. The only free variable was the harness.
Reasoning-first models (gpt-oss-120b, GLM-4.5-Air) struggled to emit usable tool calls; OpenHands-tuned Devstral produced almost nothing. For a clean comparison he anchored on qwen Q4 vs Q8.
The Code Was Almost Always Right. The Finishing Wasn't.
On the three reliable agents, correctness sat at ~0.97–0.98 median. What separated finished from failed was almost never the code:
hax wrote a correct binary in all 10 rounds — finished only 4; six hit a 100-turn cap and discarded correct work
kit declared done after 3 turns before writing tests/Makefile — no binary
pi went bimodal: one run finished but wouldn't stop (679 turns / 85M tokens); another quit after 8 turns with nothing
The False Pass
erg completed cleanly; its own make test passed green — yet the binary failed every malformed-input case. Why? Twenty-five tests encoding the same wrong reading of the spec as the code. Spec-gaming: self-verification against your own understanding is a mirror, not a test. Only an independent check against the spec (or held-out oracle) catches this.
A Bigger Model Did Not Help
Q8 (1.7× memory/latency) bought no consistent correctness gain; two agents got worse; same failure modes. Skills that say "validate inputs… run build and test before finishing" were loaded and still didn't save erg's false pass.
So What Actually Works
Environment-grounded done-gate — harness runs make build && make test and checks deliverables before allowing "done"
Independent reviewer against the spec — not the author's own tests
Graceful finalization and loop guards — commit best-effort at budget; detect no-progress loops
Notice what's not on that list: a bigger model.
Conclusion
Same brain, five bodies, one frozen judge. Nearly every failure that mattered was the body: stopping wrong, trusting the model's "done," grading its own homework. On SWE-bench, improving the scaffold moved scores more than swapping models — he watched the same pattern on a desktop bench. The model is a component. The harness is the system. Spend accordingly.