Agents can oneshot games that are actually fun. With the right guardrails, agents can execute incredibly impressive migrations in complex codebases, even rewrites in new languages. But basically all “real” software work still has human engineers driving the process. How do we get to a place where a much bigger portion of the work gets handled for us, without requiring our attention?

It seems like agents should be able to build entire software systems themselves. Why is this going so poorly in practice? What new primitives will we need to make it all work? How far can this go?

What comes after tokenmaxxing?

A lot of engineering orgs spent the first half of this year offloading as much work as possible to armies of agents and adversarial loops. The results have been pretty disappointing: mountains of dubious code, but no tsunami of incredible software. The ROI on all those tokens has been sketchy at best.