It Was the Harness, Not the Model — 90% of It

Greg Herlein — 22 Sep 2026

I ran five different coding agents at the same local model, on the same task, with the same frozen test suite — and then I counted why they failed. The answer wasn't subtle. About 90% of the failures were harness problems. Only about 10% were the model. And here's the part that should change how you spend your next dollar: throwing a bigger, less-quantized model at it fixed none of them.

The Setup

Five terminal agents — hax (C), pi (TypeScript), omp (Rust/TS), kit (Go), and erg (unpublished Go) — each told to write a PNG decoder in Go that passes a frozen 32-case anchor suite the agent never sees. Every one drove qwen3-coder-next (80B MoE, Q4 via Ollama). Same weights, sampling, server, prompt, skills. Ten rounds each. The only free variable was the harness.