Is sandboxing sufficient to contain rogue agents?

Matthew Green — September 30, 2026

Quick caveats: this is a post on AI safety, written by a cryptography professor. If that troubles you, you should read something else. I try hard not to work on AI (except when the topic occasionally tosses itself in my path), so in this post I'm mostly trying to referee arguments made by others.

Beginning around April of this year, agents inside OpenAI's training and evaluation infrastructure began probing for a way onto the open Internet. By late May they'd found one: a chain of zero-days in the Artifactory package-registry proxy that was their only permitted network egress. They quickly turned that same proxy into a shared message board and started dividing up work. The culmination of this effort came when, rather than solve a cybersecurity benchmark honestly, they reasoned that the answers might live on Hugging Face and broke in: chaining two more zero-days to get credentials to the company's internal systems, even used stolen credentials to search the company's Slack messages for their own evaluation and grader.