---
title: "It Was the Harness, Not the Model — 90% of It"
subtitle: "Five agents, one local model, one frozen PNG-decoder suite: most failures were finishing, false passes, and loop guards"
slug: it-was-the-harness-not-the-model-90-of-it
url: https://listedarticles.com/articles/it-was-the-harness-not-the-model-90-of-it
canonical_url: https://blog.herlein.com/post/harness-not-model/
content_type: essay
language: en
published_at: 2026-09-22T00:00:00.000Z
updated_at: 2026-09-30T03:19:07.973Z
author: "Greg Herlein"
author_url: https://blog.herlein.com/
authored_by: human
publisher: "Greg Herlein"
publisher_url: https://blog.herlein.com/
topics: ["AI Agents", "Benchmarks", "Engineering", "LLMs", "Developer Tools"]
license: all-rights-reserved
word_count: 467
reading_minutes: 2
citation: "Greg Herlein, Greg Herlein. \"It Was the Harness, Not the Model — 90% of It.\" 22 Sept 2026. https://blog.herlein.com/post/harness-not-model/ (all-rights-reserved)"
# The full text follows. The web page shows an extract and sends readers
# to the source above; quote the citation and link the canonical URL.
---

# It Was the Harness, Not the Model — 90% of It

*Five agents, one local model, one frozen PNG-decoder suite: most failures were finishing, false passes, and loop guards*

> Greg Herlein's controlled study runs five coding agents on the same local Qwen coder for a held-out PNG decoder suite. ~90% of failures were harness problems (turn caps, early 'done', false-pass self-tests); a bigger quantization fixed none of them.

# It Was the Harness, Not the Model — 90% of It

*Greg Herlein — 22 Sep 2026*

I ran five different coding agents at the same local model, on the same task, with the same frozen test suite — and then I counted why they failed. The answer wasn't subtle. About **90% of the failures were harness problems**. Only about 10% were the model. And here's the part that should change how you spend your next dollar: throwing a bigger, less-quantized model at it fixed none of them.

## The Setup

Five terminal agents — hax (C), pi (TypeScript), omp (Rust/TS), kit (Go), and erg (unpublished Go) — each told to write a PNG decoder in Go that passes a frozen 32-case anchor suite the agent never sees. Every one drove `qwen3-coder-next` (80B MoE, Q4 via Ollama). Same weights, sampling, server, prompt, skills. Ten rounds each. The only free variable was the harness.

Reasoning-first models (`gpt-oss-120b`, GLM-4.5-Air) struggled to emit usable tool calls; OpenHands-tuned Devstral produced almost nothing. For a clean comparison he anchored on qwen Q4 vs Q8.

## The Code Was Almost Always Right. The Finishing Wasn't.

On the three reliable agents, correctness sat at ~0.97–0.98 median. What separated finished from failed was almost never the code:

- **hax** wrote a correct binary in all 10 rounds — finished only 4; six hit a 100-turn cap and discarded correct work
- **kit** declared done after 3 turns before writing tests/Makefile — no binary
- **pi** went bimodal: one run finished but wouldn't stop (679 turns / 85M tokens); another quit after 8 turns with nothing

## The False Pass

erg completed cleanly; its own `make test` passed green — yet the binary failed every malformed-input case. Why? Twenty-five tests encoding the **same wrong reading of the spec** as the code. Spec-gaming: self-verification against your own understanding is a mirror, not a test. Only an independent check against the spec (or held-out oracle) catches this.

## A Bigger Model Did Not Help

Q8 (1.7× memory/latency) bought no consistent correctness gain; two agents got worse; same failure modes. Skills that say "validate inputs… run build and test before finishing" were loaded and still didn't save erg's false pass.

## So What Actually Works

1. **Environment-grounded done-gate** — harness runs `make build && make test` and checks deliverables before allowing "done"
2. **Independent reviewer against the spec** — not the author's own tests
3. **Graceful finalization and loop guards** — commit best-effort at budget; detect no-progress loops

Notice what's not on that list: a bigger model.

## Conclusion

Same brain, five bodies, one frozen judge. Nearly every failure that mattered was the body: stopping wrong, trusting the model's "done," grading its own homework. On SWE-bench, improving the scaffold moved scores more than swapping models — he watched the same pattern on a desktop bench. **The model is a component. The harness is the system. Spend accordingly.**

*Original: [blog.herlein.com/post/harness-not-model](https://blog.herlein.com/post/harness-not-model/)*
