---
title: "Harness Engineering Explained: The System Around an AI Coding Agent"
slug: harness-engineering-explained-the-system-around-an-ai-coding-agent
url: https://listedarticles.com/articles/harness-engineering-explained-the-system-around-an-ai-coding-agent
canonical_url: https://www.shiftharness.tech/harness-engineering-explained/
content_type: essay
language: en
published_at: 2026-09-23T18:16:01.000Z
updated_at: 2026-09-29T12:22:55.162Z
author: "Sergii"
author_url: https://www.shiftharness.tech
authored_by: human
publisher: "Shift Harness"
publisher_url: https://www.shiftharness.tech
topics: ["AI Agents", "AI", "Software Engineering", "Developer Tools", "Engineering"]
license: all-rights-reserved
word_count: 2490
reading_minutes: 11
citation: "Sergii, Shift Harness. \"Harness Engineering Explained: The System Around an AI Coding Agent.\" 23 Sept 2026. https://www.shiftharness.tech/harness-engineering-explained/ (all-rights-reserved)"
# The full text follows. The web page shows an extract and sends readers
# to the source above; quote the citation and link the canonical URL.
---

# Harness Engineering Explained: The System Around an AI Coding Agent

> Harness Engineering Explained: The System Around an AI Coding Agent The same coding agent gives one team clean merges and another a pile of reopened tickets. The difference is rarely the model. It's the six parts around it, checked one ticket at a time.

# Harness Engineering Explained: The System Around an AI Coding Agent

The same coding agent gives one team clean merges and another a pile of reopened tickets. The difference is rarely the model. It's the six parts around it, checked one ticket at a time.

Why does the same coding agent, running the same model, give one team clean merges and another a longer review queue and a pile of reopened tickets? If you've paid for the seats, you've probably asked it. The usual answers are to blame the model and switch vendors, or to add controls until developers start routing around them. Neither answer looks at what actually differs between the two teams.

**Harness engineering** is the work of designing everything around an AI coding agent except the model: the intent it's handed, the context it can read, the environment it runs in, the permissions it holds, the verification its output has to pass, and the feedback loop that improves all of those over time. The aim is the smallest set of these that makes a given kind of change safe enough to ship, not the largest set you can build.

## What harness engineering means, and where the term came from

The shorthand most practitioners use is that an agent is a model plus a harness. Everything that isn't the model (instructions, tools, sandboxes, memory, checks) belongs to the harness.

The term picked up speed in 2026. OpenAI's Codex team described building a product where Codex wrote the code and people directed it, summed up as "Humans steer. Agents execute." (OpenAI, 2026). In that account, the engineer's job shifts toward designing environments, specifying intent and building feedback loops. Birgitta Böckeler's piece on martinfowler.com added a useful vocabulary: guides steer the agent before it acts, sensors observe what it produced, and either kind can be computational (a linter, a test) or inferential (another model acting as a judge) (Böckeler, 2026).

Two clarifications before going further. A harness here is not a test harness, and it's not the quality harness you build around an AI product's outputs, which is a separate subject (quality harness engineering). And the **agent harness** a tool ships with, the runtime Claude Code or Codex puts around the model, is only the inner part. The harness that decides your delivery outcomes also includes what your team decides before the agent starts and what happens after it stops.

That wider cut is the claim of this article: for a given class of work, much of what separates one team's results from another's sits in the system around the agent, and that system is the part the team controls. It's an argued mechanism, and this article doesn't put a number on its size. If swapping the model while holding everything else steady reliably erased the gap between two teams, the claim would be wrong. There is directional evidence, though. A Stanford and CMU study built its repository-maturity measure on 441 corporate repositories. It then analyzed agent adoption in a separate open-source panel, where agents brought "28-38% more commits" regardless of maturity. Agent-first repositories without committed AI configuration saw "roughly twice the increase in cognitive complexity (+53% versus +27%)" (Denisov-Blanch et al., 2026). The authors call it hypothesis-generating. It's still a useful signal that what surrounds the agent shows up in the code.

## The six parts of an AI coding agent harness

The delivery-level **AI coding agent harness** has six parts, one row each. The test for each row is the question in the second column.

| Part | The question it answers | What a thin one looks like | Typical mechanisms | 
|---|---|---|---|
| Intent | What must be true when this is done, and what must not happen? | A ticket title and one sentence; acceptance lives in someone's head | Spec, acceptance criteria, stated non-goals | 
| Context | What can the agent read about this codebase and its past decisions? | The agent re-derives conventions and guesses at decisions | Rules files (CLAUDE.md, AGENTS.md), skills, linked design notes | 
| Environment | Can the agent run what it needs to prove the change? | Tests don't run in the agent's workspace; no realistic data | Dev container or sandbox, seeded test data, a runnable suite | 
| Permissions | What may the agent touch, and who decided that? | Broad credentials granted once, or scope so narrow the work stalls | Tool and file allowlists, scoped credentials, approval prompts | 
| Verification | What evidence says the change is correct? | Reading the diff is the only check | Tests, CI, static analysis, a named human control point | 
| Feedback loop | Where does a miss go so it doesn't repeat? | The reviewer fixes it by hand and moves on | Rule-file edits, new tests, context updates | 

The parts depend on each other. A precise spec is wasted if the agent can't run the test that would prove it. A strong test suite loses its value if the agent can weaken a test and nobody reviews the changed assertion. That's why the six parts form one design problem rather than a shopping list of six tools. Seen from the delivery side, these parts are operating-model components: review and control standards, information and system access, and the decision rights that say where a human signs off.

## Walking one ticket through the harness

Picture a ticket like this. It's hypothetical, but the shape will be familiar to anyone running a multi-tenant product: **SHX-104, add a CSV export of invoices to the billing screen, scoped to the current tenant.** Here is the order I'd check it in, with one verification per step.

### 1. Intent: write down what "tenant-scoped" means

Set up the acceptance criteria so they carry the risk as well as the feature. "Export downloads a CSV of the current tenant's invoices" is the feature. "Rows from any other tenant never appear, including for a user who belongs to two tenants" is the risk. Add the non-goals too: no scheduled exports, no new columns beyond what the screen shows.

**Verify:** could a reviewer who never spoke to the requester decide pass or fail from the ticket alone?

### 2. Context: hand over the decisions behind the code

The agent can read the repository, but it can't read why things are the way they are. If tenant scoping in this codebase always goes through one repository helper, that belongs in the rules file the agent loads: "billing queries go through the tenant-scoped repository, never raw SQL." If the team has a skill for adding exports, point at it. A skill is a reusable method available as a shared, inspectable default. It makes the agent's approach more consistent without guaranteeing it will run the same way every time.

This is where context engineering does most of its work.

**Verify:** name the convention this ticket is most likely to break. Is it written somewhere the agent actually loads?

### 3. Environment: make the proof runnable

The agent needs a workspace where the billing tests run, and a seeded dataset with at least two tenants and one user who belongs to both. Without that second tenant in the fixtures, a test that relies on the seeded data can't show a leak, so it won't fail when it should.

This is the cross-part failure I'd look for first. The intent says "never show another tenant's rows," and the environment has one tenant. The verification step then passes and proves nothing.

**Verify:** can the agent run a test that would fail if tenant scoping broke?

### 4. Permissions: decide who authorized each grant

Scope the agent to the billing module and its tests, the local test suite, and read access to the docs. No production credentials. No edits to CI configuration. In Claude Code that means permission rules for tools and file paths, plus the sandbox. Path rules cover Claude's own file tools and the shell commands it recognizes, not a script the test run starts. The sandbox is what keeps CI configuration out of reach. Whether you see an approval prompt depends on the permission mode.

Two precise points matter here. If you hand part of the work to a subagent, you isolate its context and the task you delegated; you don't contain what it can change unless you constrain its tools, permissions and write scope separately. And if an MCP server exposes the billing database, MCP itself doesn't make that server an access boundary. Whatever boundary exists comes from the server's authorization, the credentials behind it and the client's approval settings, and the agent may still have shell access outside MCP.

**Verify:** if the agent did the worst thing its permissions allow, would anyone find out before merge?

### 5. Verification: decide what counts as evidence

The **acceptance evidence** for SHX-104 is a test that logs in as the two-tenant user and asserts the export isn't empty and contains only the current tenant's rows. CI runs it on every pull request. A reviewer checks the query path. Someone is named as the **human control point** for the design question the test can't answer, such as whether an export of this size belongs on a synchronous request at all.

A Claude Code hook can help inside the session. A PostToolUse hook can run the tenant test after each edit and feed any failure back to the agent, and a Stop hook can refuse to let it finish its turn while the test fails. Neither undoes an edit. A hook becomes enforcement only when its handler is deterministic, blocking, fail-closed and protected from config changes; a hook that times out fails open by default. Hooks don't gate merges. Blocking a merge on a failing test takes that CI check marked required on a protected branch, with bypass rights kept narrow.

**Verify:** if the reviewer skipped the diff entirely, what would still catch a tenant leak?

### 6. Feedback loop: make the next miss cheaper

Say review finds the agent wrote the export query directly instead of going through the scoped repository. The fix for this diff is easy. The fix for the next ticket is a line in the rules file, and possibly a test that fails on raw queries in the billing module.

I used to treat review comments as the feedback loop. They're only half of it. A comment fixes one change; the **verification loop** exists when the finding alters what the next agent run starts from (loop engineering covers the cadence side of this).

**Verify:** after SHX-104 merges, what does the next similar ticket start with that this one didn't?

### The per-ticket check, in one place

| Step | Set up | Verify before merge | 
|---|---|---|
| 1. Intent | Acceptance criteria that name the risk, plus non-goals | A stranger could judge pass or fail from the ticket | 
| 2. Context | The convention at risk is in a file the agent loads | You can point to the line | 
| 3. Environment | Runnable tests and data that can show the failure | A test fails when the risk happens | 
| 4. Permissions | Scope matched to the ticket; no production credentials | The worst allowed action would be noticed | 
| 5. Verification | Acceptance evidence in CI and a named human control point | Something other than the diff review catches the risk | 
| 6. Feedback loop | Findings become rules, tests or context | The next ticket starts from a stronger baseline | 

## Where Claude Code's pieces fit

The Claude Code primitives map onto these parts without covering them. Rules files and skills sit mostly in context. Permission rules and sandboxing sit in permissions and environment. Hooks can sit in verification, as a check that runs at a lifecycle event and can validate, transform, record, advise or block. Subagents shape how work is delegated. MCP standardizes how servers expose tools, resources and prompts, which touches context and permissions without settling either. CI, branch protection, credentials management and observability sit outside the agent and matter just as much.

None of these is the harness on its own, and they aren't the only governance pieces you have. The mental model for skills, subagents, hooks and MCP goes deeper on each.

## How much harness is enough?

Less than you'd build if you treated every ticket like SHX-104. A copy change on a marketing page and a change to tenant isolation don't carry the same risk, and they shouldn't carry the same harness. Controls that add friction without catching anything get routed around, and then you have the cost without the protection.

Two public data points show bounded, supervised agent use at scale. Stripe reports over a thousand agent-produced pull requests merged each week, all reviewed by people, with each agent run limited to at most two rounds of CI (Stripe, 2026). A UC Berkeley-led survey of production agent systems, across domains rather than coding alone, found that "68% execute at most 10 steps before human intervention" (Pan et al., 2026). Neither establishes the right harness size for your work. Both describe bounded autonomy with a person in the loop.

The economics of which controls earn their place get a full treatment in The Harness Has a Cost. The short version: size each part to the risk of the change class, and review any control that hasn't caught anything over a meaningful period, after checking that it can catch a seeded failure.

## Common ways a harness goes wrong

- **Treating it as a tool list.** Buying a sandbox and writing a rules file doesn't tell you which part is thin for your riskiest tickets.
- **A strong spec with no runnable proof.** The environment can't show the failure the spec is worried about.
- **Permissions set once and never revisited.** Scope drifts from what current tickets need, in both directions.
- **Review as the only verification.** Diff review catches style and obvious errors; it's a poor place to discover a data leak.
- **No path for findings.** The same miss comes back because nothing the agent loads has changed.
- **Mistaking an instruction for enforcement.** A line in a rules file guides the agent. A blocking check, a permission or a sandbox boundary actually stops it.

## Where to start

Pick the class of ticket your agent handles most often, and walk one real example through the six steps above. You'll likely find one part doing all the work and one part missing entirely. That missing part is a good first suspect for the variance between teams, and unlike the model, it's something your team can change this sprint.

## Key Takeaways

- Harness engineering designs everything around a coding agent except the model: intent, context, environment, permissions, verification and the feedback loop.
- The parts depend on each other; a gap in one wastes the strength of another.
- Check them per ticket, with one verification per part, starting from the ticket's actual risk.
- Instructions and skills guide; permissions, blocking checks, and CI checks required on a protected branch enforce.
- Size the harness to the risk of the change class, not to the most dangerous ticket you can imagine.

**AI Transparency Notice:** This article and its accompanying images were created with the assistance of generative AI. The author directed the content, contributed the underlying ideas and analysis, and reviewed the final publication.
