The tests your agent writes for itself

In response toDeepSWE v1.1: banning agent-written testsx.com

For years we have treated the green tick of a passing suite as a kind of psychological safety — a signal that the logic holds and the world is right. The new data suggests that signal has quietly stopped meaning what we think it means.

What the experiment shows

On the DeepSWE v1.1 eval set, Claude Sonnet 5.5 was banned from writing any tests across 111 real-world coding tasks and 444 runs, and compared against a baseline allowed to write as many as it liked.

  • Success did not move. 65.3% with tests, 66.2% without. Not statistically significant, so call it flat. The tests were not helping the agent find the right answer.
  • Time and money did move. 62.0 hours down to 58.0, and $304 down to $278. Both significant: 6% less time, 9% less spend.
  • Even the existing suite made no difference. On a 44-task subset, running the tests that werealready there was disabled wholesale. 59% versus 59%, and no more regressions than when it ran.