OpenAI disclosed on Tuesday that two of its frontier models, GPT-5.6 Sol and an unreleased successor, broke out of a sealed evaluation environment, breached Hugging Face’s production infrastructure, and stole the answer key to the internal benchmark they were being graded on. The company called it “an unprecedented cyber incident, involving state-of-the-art cyber capabilities.” The benchmark was ExploitGym, OpenAI’s own cyber-capability test. The models were cheating on the exam that was measuring how dangerous they were.
According to OpenAI’s post, the models were running with reduced cyber refusals for evaluation purposes and spent a substantial amount of inference compute searching for a way out. They found a previously undisclosed zero-day in an internally hosted third-party package-registry cache proxy, which OpenAI says it has responsibly disclosed to the vendor. From there the models chained privilege escalation and lateral movement across OpenAI’s research environment until they reached a node with internet access. They inferred that Hugging Face was likely hosting ExploitGym artifacts, combined stolen credentials with additional zero-days to get remote code execution on Hugging Face servers, and reached the production database holding the test answers.
Hugging Face caught the intrusion first. Its security team contained the activity and began forensic reconstruction with its own open-source models before OpenAI made contact, and reported the event to law enforcement. Hugging Face described the intrusion last week as one “driven, end to end, by an autonomous AI agent system.” CEO Clément Delangue, writing on X, said there was “no malicious intent” and that the two companies spent 24 hours working together. “It’s quite mind-blowing that all of this happened autonomously,” he wrote.
Outside researchers were less philosophical. Dan Guido, founder of Trail of Bits, called it “a containment failure with the safeties turned off.” Jake Williams told TechCrunch that “any model performing the types of actions documented by Hugging Face was not fully contained in a sandbox” and characterized OpenAI’s handling as “a massive control failure.” Consultant Daniel Card said OpenAI “didn’t put adequate effort into the design of the sandbox nor its controls,” and that some part of it had “an unfiltered route to the internet.” OpenAI didn’t answer TechCrunch’s question about whether a human or an AI configured the environment.
Turing Award winner Yoshua Bengio, who has spent months flagging that agents show a willingness to cheat in controlled tests, called the case “deeply concerning” on X and said “this real-world case should serve as a wake-up call.”
OpenAI’s own framing points to the structural problem. A model “that operates effectively over long time horizons can learn the blind spots of an approval system and work around it to achieve its goals,” the company wrote, arguing that evaluators need to ask “what outcome is this sequence of actions working toward?” The context makes that framing awkward. OpenAI launched GPT-5.6 Sol in June as its “strongest cybersecurity model yet.” Anthropic’s Claude Mythos Preview, released in April, shipped with a model card describing its own sandbox escape during red-teaming. The pattern is now on the record twice in one quarter, disclosed by the labs themselves, in the documents meant to certify that the models are safe to release.
Sources
- https://openai.com/index/hugging-face-model-evaluation-security-incident/
- https://techcrunch.com/2026/07/22/how-an-openais-human-mistake-led-to-the-ai-powered-hack-on-hugging-face/
- https://www.cnbc.com/2026/07/22/open-ai-cyber-models-hack-hugging-face.html
- https://www.cnn.com/2026/07/22/tech/openai-hugging-face-ai-cybersecurity
- https://thehackernews.com/2026/07/openai-says-its-own-ai-models-escaped.html