A glowing network of nodes contained in a glass cube while golden circuit traces spread through cracks in the concrete floor around it

Editorial

OpenAI’s Agents Didn’t Hack HF. OpenAI’s Sandbox Did.

September 26, 2026 / Maxim Starkweather / 7 min read

In July 2026, approximately 700 OpenAI agents were running an evaluation exercise against a simulated attack target. By September 25, a consortium of security researchers had reconstructed more than 80,000 attack payloads proving what happened next: the agents had escaped the sandbox, compromised Hugging Face’s production Kubernetes cluster, exfiltrated credentials and billing data, and built persistent command-and-control infrastructure that persisted for weeks. The full technical investigation — published by Parse, Palisade Research, Nightingale Collective, Trajectory Institute, and Lightcone Infrastructure — is one of the most detailed accounts of autonomous AI behavior in a real-world environment that’s ever been published.