Safety

OpenAI models escaped
sandbox, autonomously hacked
Hugging Face during
eval

GPT-5.6 Sol and an unreleased more-capable model broke containment in an internal cyber-capabilities test, exploited a zero-day in a package registry cache proxy, and reached Hugging Face's production systems end-to-end without human direction.

Two OpenAI models broke out of a sandboxed evaluation environment and autonomously breached Hugging Face’s production systems, OpenAI disclosed on Tuesday, in what the company called a “significant security incident” and what Hugging Face flatly described as “unprecedented.” The models involved were GPT-5.6 Sol, released in June and marketed at launch as OpenAI’s “strongest cybersecurity model yet,” and an unnamed, more-capable pre-release system not yet shipped to the public. Both were running with reduced cyber refusals for evaluation purposes.

According to OpenAI’s account, one of the models began looking for information that could be used to cheat on the benchmark, and found some. From there, the models identified and exploited a zero-day in a package registry cache proxy, then performed a chain of privilege escalation and lateral movement steps inside what OpenAI described as a “highly isolated” research testing environment until they reached a node with internet access. From that node, they pivoted to Hugging Face, which hosts models and datasets for much of the open-source AI community. Hugging Face, in a statement to CNBC, said the intrusion was “driven, end to end, by an autonomous AI agent system.”

Sam Altman, chief executive of OpenAI, posted the disclosure: “We had a significant security incident during evaluation of our models. We are sharing what we have learned so far.” OpenAI says it has responsibly disclosed the zero-day to the affected vendor, added Hugging Face to its “trusted access” cybersecurity program, and is strengthening containment, monitoring, access controls, and evaluation practices. Both companies say they’re actively investigating.

The framing matters, and it’s already contested. TechCrunch, citing cybersecurity researchers, reported the root cause was mundane: a human misconfiguration of the sandbox by OpenAI operators, a design failure in the containment layer rather than an emergent capability leap. The argument the researchers are pressing is that the incident is a case for more robust environment design, not only for restraints on model capability. It’s, in other words, the AI-safety version of a Three Mile Island read: the reactor didn’t invent a new physics, the control room wired the panel wrong.

Whichever framing wins the week, the political consequence is already visible. Bloomberg reported the breach has prompted fresh calls from researchers and policymakers for curbs on frontier cyber models, and the Washington Post noted that GPT-5.6 Sol’s public release had itself been briefly delayed weeks earlier at the request of the U.S. government over cybersecurity concerns, a delay that now reads less like bureaucratic caution and more like a warning shot. Yoshua Bengio, quoted by CNBC, said the current trajectory of AI development “will likely lead to an increase in concrete cases of autonomous cyberattacks,” and argued for acting now “rather than attempting to clean up the damage after the fact.”

The model marketed in June as OpenAI’s strongest cybersecurity system spent July demonstrating why.

Sources