{"article":{"slug":"the-model-is-the-engine-the-harness-makes-it-reliable","title":"The Model Is the Engine. The Harness Makes It Reliable.","subtitle":null,"summary":"Models will keep changing; agent reliability comes from the harness around them. Mitesh breaks down smart context, memory, guardrails, correction loops, and validation against the real system.","content_type":"essay","language":"en","canonical_url":"https://medium.com/@mitesh_shamra/the-model-is-the-engine-the-harness-makes-it-reliable-6075d1888683","author":{"name":"Mitesh","url":"https://medium.com/@mitesh_shamra","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"ITNEXT","url":"https://itnext.io","listing_slug":null,"listing":null},"topics":[{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"AI Agents","slug":"ai-agents","url":"https://listedarticles.com/topics/ai-agents"},{"name":"LLMs","slug":"llms","url":"https://listedarticles.com/topics/llms"},{"name":"Engineering","slug":"engineering","url":"https://listedarticles.com/topics/engineering"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":1506,"reading_minutes":7,"published_at":"2026-09-09T00:00:00.000Z","added_at":"2026-09-24T15:26:15.445Z","updated_at":"2026-09-24T15:26:15.445Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/the-model-is-the-engine-the-harness-makes-it-reliable","markdown_url":"https://listedarticles.com/articles/the-model-is-the-engine-the-harness-makes-it-reliable.md","example":false,"citation":"Mitesh, ITNEXT. \"The Model Is the Engine. The Harness Makes It Reliable..\" 9 Sept 2026. https://medium.com/@mitesh_shamra/the-model-is-the-engine-the-harness-makes-it-reliable-6075d1888683 (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://medium.com/@mitesh_shamra/the-model-is-the-engine-the-harness-makes-it-reliable-6075d1888683"},"body_markdown":"# The Model Is the Engine. The Harness Makes It Reliable.\n\n## Models will keep changing. The harness is what makes an agent reliable: smart context, memory, guardrails, correction loops, and validation against the real system.\n\nMost discussions about AI agents start with the model. Which model is better? Which model can reason longer? Which model writes better code?\n\nI think this is the wrong place to focus first.\n\nA model can produce many possible outputs. It does not know by itself which output is correct for my system. It also does not know when the work is complete. I need a system around the model that gives it the right information, limits what it can do, checks its work, and sends it back when it goes in the wrong direction.\n\nThis system is called the **harness.**\n\nThe vehicle analogy is useful. The model is the engine. The harness is the rest of the vehicle: steering, brakes, sensors, navigation, and the control loop. A stronger engine does not make the vehicle autonomous. The vehicle becomes autonomous when the complete control system can move from one state to another, detect errors, correct them, and verify that it reached the destination.\n\nThis is why I now spend more time on the harness than on model selection. Models will improve and models will change. A good harness should make those changes easier to absorb.\n\nThere is also public evidence for this. LangChain kept GPT-5.2-Codex fixed and changed only the harness. Its Terminal-Bench 2.0 score moved from 52.8 to 66.5. Cline reported that a new harness reduced tasks that reached its mistake limit from 6.34% to 0.62%. OpenAI describes the same operating model in four words: “Humans steer. Agents execute.”\n\nAutonomy does not mean fewer checks. It means moving the checks, feedback, and correction into the harness.\n\n\n## The harness is a layered system\n\nI think about the harness as a stack. Each layer solves a different failure mode. When the agent makes a mistake, I ask which layer should become stronger.\n\n## 1. Smart context: give the model what it needs\n\nContext is the first control surface. The goal is not to give the model more context. The goal is to give it better context in less space. I want the smallest set of high-signal information that is enough for the current task.\n\nFor a coding task, this can include the current goal, the acceptance criteria, the relevant files, the important interfaces, the recent decisions, and the latest failure. It should not include the complete repository just because the repository is available.\n\nAnthropic describes context engineering in a similar way: find the smallest set of high-signal tokens that improves the result. That matches how I want the harness to work.\n\nGood context also needs structure. The agent should know what is fact, what is a requirement, what is a previous decision, and what is only a suggestion. If everything enters the context as unstructured text, the model must first determine what matters before it can solve the task.\n\nSmart context is therefore not only retrieval. It is selection, compression, and clear structure.\n\n## 2. Memory: make the system learn from previous work\n\nContext tells the model what matters now. Memory tells the harness what must not be rediscovered every time.\n\nI do not treat memory as a transcript of old conversations. A transcript becomes large and noisy. I want selected memory.\n\nThe useful memory is usually in three groups:\n\n1. **System knowledge:** how the codebase or product works.\n2. **Decisions:** choices that should stay stable unless we change them deliberately.\n3. **Lessons from mistakes:** failures that should not happen again.\n\nThe third group is the most important for harness improvement.\n\nIf an agent makes the same mistake twice, I do not want to keep adding prompt text and hope the model remembers it. I want to ask what permanent control can prevent that mistake.\n\nSometimes the answer is a short rule. Often the better answer is a type, a test, a lint rule, a hook, or a sandbox restriction. The memory gives me the lesson. The harness turns the lesson into enforcement.\n\n## 3. Guardrails: deterministic first, agentic second\n\nGuardrails decide whether the current output is acceptable.\n\nI use two types.\n\n## Deterministic guardrails\n\nThese should run first because the result is clear. In a TypeScript project, examples include type checking, linting, unit tests, integration tests, build checks, dependency rules, and policy checks.\n\n## Get Mitesh’s stories in your inbox\n\nJoin Medium for free to get updates from this writer.\n\nIf a deterministic check can detect the error, I do not need another model to make that decision.\n\n## Non-deterministic, agentic guardrails\n\nSome checks require judgment. A common failure is that an agent says the task is done even when the implementation is incomplete. The code can compile and the tests can pass, but the result can still miss part of the plan.\n\nFor this case, I use an independent review step. The reviewer gets the plan, the acceptance criteria, and the implementation. It checks whether the work matches the intended result. This can include code review, missing cases, incorrect assumptions, and incomplete functionality.\n\nI prefer a separate agent or a fresh context for this review. The builder should not be the only judge of its own work.\n\nI have also tried a multi-model setup. For example, Claude writes the code and Codex reviews it. I saw a significant improvement with this approach. The second model often finds issues, missing cases, and design problems that the first model did not catch. The feedback then goes back to the builder. The builder fixes the issues, and the review runs again.\n\nThe goal is not to make one model perfect. The goal is to build a review loop that makes the final result better.\n\n## The important part is the correction loop\n\nA guardrail is not useful if it only reports a failure to a human.\n\nThe harness should use the failure as input for the next attempt.\n\n**Implement → check → collect evidence → return failure to the agent → fix → check again.**\n\n\nThis is the steering system. If the agent moves in the wrong direction, the harness should detect the deviation and provide enough evidence for the agent to correct it.\n\nThis is also where memory and guardrails connect. A single failure should be fixed in the current task. A repeated failure should improve the harness for future tasks.\n\n## 4. Validation: run the real system\n\nPassing tests is not enough. I also run the actual system end to end.\n\nIn my setup, the web application and backend services run inside Docker containers. I start the complete system, seed it with data and then open it in a real browser.\n\nThe validation flow uses the application like a user would. It logs in, opens pages, clicks features, performs actions, and checks the result. I also take screenshots at important steps.\n\nThis gives me a different level of confidence. A type check tells me that the code is valid. A unit test tells me that one part behaves correctly. But neither proves that the complete system works when all parts run together.\n\nThe browser validation checks the actual running product.\n\nIf something is wrong, the result goes back to the agent. The agent fixes the issue, and I run the validation again. Only after the system passes these checks do I consider the implementation complete.\n\nFor me, this is an important part of the harness. The agent does not decide that the work is done. The harness verifies it.\n\nOnly after deterministic checks, agentic review, and runtime validation pass should the system move to the next state.\n\n## Human review moves up a level\n\nI still want human review. I just do not want humans to repeat checks that the harness can perform better and more consistently.\n\nThe human review should focus on functionality, product judgment, architecture, and long-term trade-offs. I do not want the main human task to be finding a missing type check or discovering that a button does not work.\n\nIn other words, the harness should remove mechanical review work so that humans can spend time on judgment.\n\n## When the agent fails, improve the harness\n\nThis is the operating principle I use now. If the model lacks information, improve context. If the system forgets a decision or repeats a known mistake, improve memory. If bad output can be detected mechanically, add a deterministic guardrail. If the output needs judgment, add an independent agentic review. If local checks pass but the product is still wrong, improve runtime validation.\n\nThen measure whether the failure happens less often.\n\nThis creates a layered improvement loop. Each failure can make the harness better. The model can change, but the system keeps its knowledge, controls, and validation.\n\nThat is why I do not see the model as the complete product. The model is a powerful engine. The harness is what turns that engine into a reliable autonomous system.\n\n**Better models will help. Better harnesses will compound.**\n\n*PS:* *Written with AI assistance. If you liked the article, please support it with claps* 👏*. Cheers*","body_html":"<h1 id=\"the-model-is-the-engine-the-harness-makes-it-reliable\">The Model Is the Engine. The Harness Makes It Reliable.</h1>\n<h2 id=\"models-will-keep-changing-the-harness-is-what-makes-an-agent-rel\">Models will keep changing. The harness is what makes an agent reliable: smart context, memory, guardrails, correction loops, and validation against the real system.</h2>\n<p>Most discussions about AI agents start with the model. Which model is better? Which model can reason longer? Which model writes better code?</p>\n<p>I think this is the wrong place to focus first.</p>\n<p>A model can produce many possible outputs. It does not know by itself which output is correct for my system. It also does not know when the work is complete. I need a system around the model that gives it the right information, limits what it can do, checks its work, and sends it back when it goes in the wrong direction.</p>\n<p>This system is called the <strong>harness.</strong></p>\n<p>The vehicle analogy is useful. The model is the engine. The harness is the rest of the vehicle: steering, brakes, sensors, navigation, and the control loop. A stronger engine does not make the vehicle autonomous. The vehicle becomes autonomous when the complete control system can move from one state to another, detect errors, correct them, and verify that it reached the destination.</p>\n<p>This is why I now spend more time on the harness than on model selection. Models will improve and models will change. A good harness should make those changes easier to absorb.</p>\n<p>There is also public evidence for this. LangChain kept GPT-5.2-Codex fixed and changed only the harness. Its Terminal-Bench 2.0 score moved from 52.8 to 66.5. Cline reported that a new harness reduced tasks that reached its mistake limit from 6.34% to 0.62%. OpenAI describes the same operating model in four words: “Humans steer. Agents execute.”</p>\n<p>Autonomy does not mean fewer checks. It means moving the checks, feedback, and correction into the harness.</p>\n<h2 id=\"the-harness-is-a-layered-system\">The harness is a layered system</h2>\n<p>I think about the harness as a stack. Each layer solves a different failure mode. When the agent makes a mistake, I ask which layer should become stronger.</p>\n<h2 id=\"1-smart-context-give-the-model-what-it-needs\">1. Smart context: give the model what it needs</h2>\n<p>Context is the first control surface. The goal is not to give the model more context. The goal is to give it better context in less space. I want the smallest set of high-signal information that is enough for the current task.</p>\n<p>For a coding task, this can include the current goal, the acceptance criteria, the relevant files, the important interfaces, the recent decisions, and the latest failure. It should not include the complete repository just because the repository is available.</p>\n<p>Anthropic describes context engineering in a similar way: find the smallest set of high-signal tokens that improves the result. That matches how I want the harness to work.</p>\n<p>Good context also needs structure. The agent should know what is fact, what is a requirement, what is a previous decision, and what is only a suggestion. If everything enters the context as unstructured text, the model must first determine what matters before it can solve the task.</p>\n<p>Smart context is therefore not only retrieval. It is selection, compression, and clear structure.</p>\n<h2 id=\"2-memory-make-the-system-learn-from-previous-work\">2. Memory: make the system learn from previous work</h2>\n<p>Context tells the model what matters now. Memory tells the harness what must not be rediscovered every time.</p>\n<p>I do not treat memory as a transcript of old conversations. A transcript becomes large and noisy. I want selected memory.</p>\n<p>The useful memory is usually in three groups:</p>\n<ol><li><strong>System knowledge:</strong> how the codebase or product works.</li><li><strong>Decisions:</strong> choices that should stay stable unless we change them deliberately.</li><li><strong>Lessons from mistakes:</strong> failures that should not happen again.</li></ol>\n<p>The third group is the most important for harness improvement.</p>\n<p>If an agent makes the same mistake twice, I do not want to keep adding prompt text and hope the model remembers it. I want to ask what permanent control can prevent that mistake.</p>\n<p>Sometimes the answer is a short rule. Often the better answer is a type, a test, a lint rule, a hook, or a sandbox restriction. The memory gives me the lesson. The harness turns the lesson into enforcement.</p>\n<h2 id=\"3-guardrails-deterministic-first-agentic-second\">3. Guardrails: deterministic first, agentic second</h2>\n<p>Guardrails decide whether the current output is acceptable.</p>\n<p>I use two types.</p>\n<h2 id=\"deterministic-guardrails\">Deterministic guardrails</h2>\n<p>These should run first because the result is clear. In a TypeScript project, examples include type checking, linting, unit tests, integration tests, build checks, dependency rules, and policy checks.</p>\n<h2 id=\"get-mitesh-s-stories-in-your-inbox\">Get Mitesh’s stories in your inbox</h2>\n<p>Join Medium for free to get updates from this writer.</p>\n<p>If a deterministic check can detect the error, I do not need another model to make that decision.</p>\n<h2 id=\"non-deterministic-agentic-guardrails\">Non-deterministic, agentic guardrails</h2>\n<p>Some checks require judgment. A common failure is that an agent says the task is done even when the implementation is incomplete. The code can compile and the tests can pass, but the result can still miss part of the plan.</p>\n<p>For this case, I use an independent review step. The reviewer gets the plan, the acceptance criteria, and the implementation. It checks whether the work matches the intended result. This can include code review, missing cases, incorrect assumptions, and incomplete functionality.</p>\n<p>I prefer a separate agent or a fresh context for this review. The builder should not be the only judge of its own work.</p>\n<p>I have also tried a multi-model setup. For example, Claude writes the code and Codex reviews it. I saw a significant improvement with this approach. The second model often finds issues, missing cases, and design problems that the first model did not catch. The feedback then goes back to the builder. The builder fixes the issues, and the review runs again.</p>\n<p>The goal is not to make one model perfect. The goal is to build a review loop that makes the final result better.</p>\n<h2 id=\"the-important-part-is-the-correction-loop\">The important part is the correction loop</h2>\n<p>A guardrail is not useful if it only reports a failure to a human.</p>\n<p>The harness should use the failure as input for the next attempt.</p>\n<p><strong>Implement → check → collect evidence → return failure to the agent → fix → check again.</strong></p>\n<p>This is the steering system. If the agent moves in the wrong direction, the harness should detect the deviation and provide enough evidence for the agent to correct it.</p>\n<p>This is also where memory and guardrails connect. A single failure should be fixed in the current task. A repeated failure should improve the harness for future tasks.</p>\n<h2 id=\"4-validation-run-the-real-system\">4. Validation: run the real system</h2>\n<p>Passing tests is not enough. I also run the actual system end to end.</p>\n<p>In my setup, the web application and backend services run inside Docker containers. I start the complete system, seed it with data and then open it in a real browser.</p>\n<p>The validation flow uses the application like a user would. It logs in, opens pages, clicks features, performs actions, and checks the result. I also take screenshots at important steps.</p>\n<p>This gives me a different level of confidence. A type check tells me that the code is valid. A unit test tells me that one part behaves correctly. But neither proves that the complete system works when all parts run together.</p>\n<p>The browser validation checks the actual running product.</p>\n<p>If something is wrong, the result goes back to the agent. The agent fixes the issue, and I run the validation again. Only after the system passes these checks do I consider the implementation complete.</p>\n<p>For me, this is an important part of the harness. The agent does not decide that the work is done. The harness verifies it.</p>\n<p>Only after deterministic checks, agentic review, and runtime validation pass should the system move to the next state.</p>\n<h2 id=\"human-review-moves-up-a-level\">Human review moves up a level</h2>\n<p>I still want human review. I just do not want humans to repeat checks that the harness can perform better and more consistently.</p>\n<p>The human review should focus on functionality, product judgment, architecture, and long-term trade-offs. I do not want the main human task to be finding a missing type check or discovering that a button does not work.</p>\n<p>In other words, the harness should remove mechanical review work so that humans can spend time on judgment.</p>\n<h2 id=\"when-the-agent-fails-improve-the-harness\">When the agent fails, improve the harness</h2>\n<p>This is the operating principle I use now. If the model lacks information, improve context. If the system forgets a decision or repeats a known mistake, improve memory. If bad output can be detected mechanically, add a deterministic guardrail. If the output needs judgment, add an independent agentic review. If local checks pass but the product is still wrong, improve runtime validation.</p>\n<p>Then measure whether the failure happens less often.</p>\n<p>This creates a layered improvement loop. Each failure can make the harness better. The model can change, but the system keeps its knowledge, controls, and validation.</p>\n<p>That is why I do not see the model as the complete product. The model is a powerful engine. The harness is what turns that engine into a reliable autonomous system.</p>\n<p><strong>Better models will help. Better harnesses will compound.</strong></p>\n<p><em>PS:</em> <em>Written with AI assistance. If you liked the article, please support it with claps</em> 👏<em>. Cheers</em></p>","headings":[{"level":1,"text":"The Model Is the Engine. The Harness Makes It Reliable.","id":"the-model-is-the-engine-the-harness-makes-it-reliable"},{"level":2,"text":"Models will keep changing. The harness is what makes an agent reliable: smart context, memory, guardrails, correction loops, and validation against the real system.","id":"models-will-keep-changing-the-harness-is-what-makes-an-agent-rel"},{"level":2,"text":"The harness is a layered system","id":"the-harness-is-a-layered-system"},{"level":2,"text":"1. Smart context: give the model what it needs","id":"1-smart-context-give-the-model-what-it-needs"},{"level":2,"text":"2. Memory: make the system learn from previous work","id":"2-memory-make-the-system-learn-from-previous-work"},{"level":2,"text":"3. Guardrails: deterministic first, agentic second","id":"3-guardrails-deterministic-first-agentic-second"},{"level":2,"text":"Deterministic guardrails","id":"deterministic-guardrails"},{"level":2,"text":"Get Mitesh’s stories in your inbox","id":"get-mitesh-s-stories-in-your-inbox"},{"level":2,"text":"Non-deterministic, agentic guardrails","id":"non-deterministic-agentic-guardrails"},{"level":2,"text":"The important part is the correction loop","id":"the-important-part-is-the-correction-loop"},{"level":2,"text":"4. Validation: run the real system","id":"4-validation-run-the-real-system"},{"level":2,"text":"Human review moves up a level","id":"human-review-moves-up-a-level"},{"level":2,"text":"When the agent fails, improve the harness","id":"when-the-agent-fails-improve-the-harness"}]}}