An agent can finish a change in minutes and leave it waiting for review for two days. The generation was fast. The delivery wasn’t.

That’s the problem I keep coming back to as I read this month’s AI announcements. Models are getting cheaper. Platforms are taking over more of the work required to run agents. But someone still has to establish whether the result is correct and belongs in the system.

For a team already producing more changes than it can review, cheaper generation adds to the backlog. The constraint has moved, and the way we measure progress needs to move with it.

On September 22, Anthropic released Claude Opus 5.5 at $4 per million input tokens and $20 per million output tokens. Anthropic reports roughly 40% lower cost on typical workloads than Opus 5, accounting for efficiency improvements. That’s the vendor’s finding; your workload still needs its own evaluation. Anthropic’s announcement.