{"article":{"slug":"the-tests-your-agent-writes-for-itself","title":"The tests your agent writes for itself","subtitle":null,"summary":"Mashhood Rastgar responds to Kun Chen's DeepSWE v1.1 experiment, where banning Claude Sonnet 5.5 from writing its own tests left success flat while cutting time and cost, arguing the real lesson is about unrequested agent-written tests versus human-specified tests that encode intent.","content_type":"essay","language":"en","canonical_url":"https://karachiwala.dev/writing/the-tests-your-agent-writes-for-itself","author":{"name":"Mashhood Rastgar","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"karachiwala.dev","url":"https://karachiwala.dev/","listing_slug":null,"listing":null},"topics":[{"name":"AI Agents","slug":"ai-agents","url":"https://listedarticles.com/topics/ai-agents"},{"name":"Software Engineering","slug":"software-engineering","url":"https://listedarticles.com/topics/software-engineering"},{"name":"Testing","slug":"testing","url":"https://listedarticles.com/topics/testing"}],"about_listings":[],"cover_image_url":"https://karachiwala.dev/og/the-tests-your-agent-writes-for-itself.jpg","license":"all-rights-reserved","word_count":768,"reading_minutes":3,"published_at":"2026-10-08T00:00:00.000Z","added_at":"2026-10-09T08:16:55.802Z","updated_at":"2026-10-09T08:16:55.802Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":false},"profile_url":"https://listedarticles.com/articles/the-tests-your-agent-writes-for-itself","markdown_url":"https://listedarticles.com/articles/the-tests-your-agent-writes-for-itself.md","example":false,"citation":"Mashhood Rastgar, karachiwala.dev. \"The tests your agent writes for itself.\" 8 Oct 2026. https://karachiwala.dev/writing/the-tests-your-agent-writes-for-itself (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://karachiwala.dev/writing/the-tests-your-agent-writes-for-itself"},"body_markdown":"# The tests your agent writes for itself\n\nIn response to[DeepSWE v1.1: banning agent-written tests](https://x.com/kunchenguid/status/2108030810691629403)x.com\n\nFor years we have treated the green tick of a passing suite as a kind of psychological safety — a signal that the logic holds and the world is right. The new data suggests that signal has quietly stopped meaning what we think it means.\n\n## What the experiment shows\n\nOn the DeepSWE v1.1 eval set, Claude Sonnet 5.5 was banned from writing any tests across 111 real-world coding tasks and 444 runs, and compared against a baseline allowed to write as many as it liked.\n\n- **Success did not move.** 65.3% with tests, 66.2% without. Not statistically significant, so call it flat. The tests were not helping the agent find the right answer.\n- **Time and money did move.** 62.0 hours down to 58.0, and $304 down to $278. Both significant: 6% less time, 9% less spend.\n- **Even the existing suite made no difference.** On a 44-task subset, running the tests that were*already there* was disabled wholesale. 59% versus 59%, and no more regressions than when it ran.\n\nTaking the tests away cost nothing, and saved something.\n\n## The case for taking it seriously\n\n**The agent is marking its own homework.** The implementation and the tests are both the model's interpretation of your intent, born from the same context window. A test written from the same understanding as the code is not independent evidence about that code. If the understanding was wrong, you get a wrong implementation and a green suite confirming its own error. That is a closed loop, not a check.\n\n**It is not a finding about tests.** It is a finding about tests *nobody asked for* — the ones written unprompted, as a reflex, because you requested a change. Bun's rewrite leaned hard on its suite, and that suite was a human asset curated over years to encode intended behaviour. Nothing in this eval touches it. The comparison is tests a human specified against tests a model volunteered.\n\n**Stop blaming the codebase.** These runs are not on legacy spaghetti. They are on repositories like FastAPI. If the theory is that agents write bad tests because the surrounding code is bad, that theory has to explain FastAPI.\n\n## The case against\n\n**\"Tests are for regressions, and a single-task benchmark cannot see them.\"** The strongest objection, and already answered. The 44-task subset is precisely that experiment, and disabling the suite produced no extra regressions.\n\n**\"Then tests never worked.\"** Too strong, and the data does not say it. TDD was genuinely useful to humans for years during first-pass development. The finding is narrower and stranger: agents do not appear to inherit the benefit humans got from it. That is worth investigating, not celebrating.\n\n**\"Mutation testing fixes it.\"** Half right. Mutation testing cannot tell an intended change from an unintended one, so when a test nobody asked for fails, it cannot say whether the code broke the test or the test was wrong. It is not a source of truth. It is still a good **eval**: break the implementation deliberately and see whether anything notices. A suite that survives deliberate breakage is provably decoration. Use it to delete, not to trust.\n\n**\"This settles testing.\"** It does not. The agents wrote 17 end-to-end tests against more than 3,000 unit and integration tests. Nothing here says anything about that layer, in either direction.\n\n## Where that leaves me\n\nA test is worth exactly what the intent behind it is worth. It is an intermediate representation of something a person decided. When the agent writes both the code and the check, the chain becomes a loop: two artefacts expressing one guess, and a build that goes green when they agree with each other.\n\nFor anyone in the trenches with these tools today:\n\n- **Stop accepting tests you did not ask for.** If the agent volunteered it, it is not evidence.\n- **Specify the cases yourself** where you can express intent better as a test than as a sentence of requirements.\n- **Run mutation testing to delete** , not to trust. Find the dead weight and remove it.\n- **Read the diff on high-risk paths.** The suite was written by the same thing that wrote the code.\n\nAnd if the cost of being wrong is an afternoon, just code away. Most work does not need the ceremony, and pretending otherwise is how the ceremony got so expensive.\n\n## References\n\n- [Kun Chen: AI-written unit and integration tests are empirically unhelpful](https://x.com/kunchenguid/status/2108030810691629403) x.com\n- [Kun Chen: answering the pushback, and what a source of truth has to be](https://x.com/kunchenguid/status/2108244512808243470) x.com\n- [Julian Harris: a test protects against side effects of previously written code](https://x.com/julianharris/status/2108111666785198129) x.com","body_html":"<h1 id=\"the-tests-your-agent-writes-for-itself\">The tests your agent writes for itself</h1>\n<p>In response to<a href=\"https://x.com/kunchenguid/status/2108030810691629403\" rel=\"nofollow ugc noopener\">DeepSWE v1.1: banning agent-written tests</a>x.com</p>\n<p>For years we have treated the green tick of a passing suite as a kind of psychological safety — a signal that the logic holds and the world is right. The new data suggests that signal has quietly stopped meaning what we think it means.</p>\n<h2 id=\"what-the-experiment-shows\">What the experiment shows</h2>\n<p>On the DeepSWE v1.1 eval set, Claude Sonnet 5.5 was banned from writing any tests across 111 real-world coding tasks and 444 runs, and compared against a baseline allowed to write as many as it liked.</p>\n<ul><li><strong>Success did not move.</strong> 65.3% with tests, 66.2% without. Not statistically significant, so call it flat. The tests were not helping the agent find the right answer.</li><li><strong>Time and money did move.</strong> 62.0 hours down to 58.0, and $304 down to $278. Both significant: 6% less time, 9% less spend.</li><li><strong>Even the existing suite made no difference.</strong> On a 44-task subset, running the tests that were<em>already there</em> was disabled wholesale. 59% versus 59%, and no more regressions than when it ran.</li></ul>\n<p>Taking the tests away cost nothing, and saved something.</p>\n<h2 id=\"the-case-for-taking-it-seriously\">The case for taking it seriously</h2>\n<p><strong>The agent is marking its own homework.</strong> The implementation and the tests are both the model&#39;s interpretation of your intent, born from the same context window. A test written from the same understanding as the code is not independent evidence about that code. If the understanding was wrong, you get a wrong implementation and a green suite confirming its own error. That is a closed loop, not a check.</p>\n<p><strong>It is not a finding about tests.</strong> It is a finding about tests <em>nobody asked for</em> — the ones written unprompted, as a reflex, because you requested a change. Bun&#39;s rewrite leaned hard on its suite, and that suite was a human asset curated over years to encode intended behaviour. Nothing in this eval touches it. The comparison is tests a human specified against tests a model volunteered.</p>\n<p><strong>Stop blaming the codebase.</strong> These runs are not on legacy spaghetti. They are on repositories like FastAPI. If the theory is that agents write bad tests because the surrounding code is bad, that theory has to explain FastAPI.</p>\n<h2 id=\"the-case-against\">The case against</h2>\n<p><strong>&quot;Tests are for regressions, and a single-task benchmark cannot see them.&quot;</strong> The strongest objection, and already answered. The 44-task subset is precisely that experiment, and disabling the suite produced no extra regressions.</p>\n<p><strong>&quot;Then tests never worked.&quot;</strong> Too strong, and the data does not say it. TDD was genuinely useful to humans for years during first-pass development. The finding is narrower and stranger: agents do not appear to inherit the benefit humans got from it. That is worth investigating, not celebrating.</p>\n<p><strong>&quot;Mutation testing fixes it.&quot;</strong> Half right. Mutation testing cannot tell an intended change from an unintended one, so when a test nobody asked for fails, it cannot say whether the code broke the test or the test was wrong. It is not a source of truth. It is still a good <strong>eval</strong>: break the implementation deliberately and see whether anything notices. A suite that survives deliberate breakage is provably decoration. Use it to delete, not to trust.</p>\n<p><strong>&quot;This settles testing.&quot;</strong> It does not. The agents wrote 17 end-to-end tests against more than 3,000 unit and integration tests. Nothing here says anything about that layer, in either direction.</p>\n<h2 id=\"where-that-leaves-me\">Where that leaves me</h2>\n<p>A test is worth exactly what the intent behind it is worth. It is an intermediate representation of something a person decided. When the agent writes both the code and the check, the chain becomes a loop: two artefacts expressing one guess, and a build that goes green when they agree with each other.</p>\n<p>For anyone in the trenches with these tools today:</p>\n<ul><li><strong>Stop accepting tests you did not ask for.</strong> If the agent volunteered it, it is not evidence.</li><li><strong>Specify the cases yourself</strong> where you can express intent better as a test than as a sentence of requirements.</li><li><strong>Run mutation testing to delete</strong> , not to trust. Find the dead weight and remove it.</li><li><strong>Read the diff on high-risk paths.</strong> The suite was written by the same thing that wrote the code.</li></ul>\n<p>And if the cost of being wrong is an afternoon, just code away. Most work does not need the ceremony, and pretending otherwise is how the ceremony got so expensive.</p>\n<h2 id=\"references\">References</h2>\n<ul><li><a href=\"https://x.com/kunchenguid/status/2108030810691629403\" rel=\"nofollow ugc noopener\">Kun Chen: AI-written unit and integration tests are empirically unhelpful</a> x.com</li><li><a href=\"https://x.com/kunchenguid/status/2108244512808243470\" rel=\"nofollow ugc noopener\">Kun Chen: answering the pushback, and what a source of truth has to be</a> x.com</li><li><a href=\"https://x.com/julianharris/status/2108111666785198129\" rel=\"nofollow ugc noopener\">Julian Harris: a test protects against side effects of previously written code</a> x.com</li></ul>","headings":[{"level":1,"text":"The tests your agent writes for itself","id":"the-tests-your-agent-writes-for-itself"},{"level":2,"text":"What the experiment shows","id":"what-the-experiment-shows"},{"level":2,"text":"The case for taking it seriously","id":"the-case-for-taking-it-seriously"},{"level":2,"text":"The case against","id":"the-case-against"},{"level":2,"text":"Where that leaves me","id":"where-that-leaves-me"},{"level":2,"text":"References","id":"references"}]}}