Safety

OpenAI Can't Rule
Out 'Critical' Cyber
Capability in Astra,
Pauses Work

The company said Friday that internal evaluations of its unreleased Astra model showed offensive cyber performance strong enough to trigger the top tier of its Preparedness Framework — halting some internal development, isolating testing, and inviting government agencies to probe the model.

OpenAI said Friday it can’t rule out that Astra, an unreleased model still in evaluation, has crossed into the “Critical” cyber tier of the company’s Preparedness Framework, the highest designation in a rubric OpenAI first published in December 2023. It’s the first time the company has flagged a model at that threshold.

“While we continue to benchmark and assess this model, our preliminary evaluations indicate strong enough performance that we cannot rule out Critical capability level at this time,” OpenAI wrote in a disclosure first reported by Axios.

The framework defines Critical by two criteria, either of which is enough to trip it. Criterion 1: a model that can identify and develop functional zero-day exploits across many hardened real-world systems without human intervention. Criterion 2: a model that can devise and execute end-to-end novel cyberattacks against hardened targets given only a high-level goal. Both describe autonomous offensive capability, not merely assisted exploitation, a distinction Bloomberg and Reuters emphasized in their coverage.

OpenAI’s response was structural. The company paused internal work that doesn’t yet meet strengthened security requirements, moved Astra’s development into isolated testing environments, and deployed universal monitoring across agentic applications during training and evaluation, with monitors reading chain-of-thought and interrupting high-risk activity. It will share recommended controls with third-party testers and invited government agencies and select AI safety organizations to probe the model directly.

The timing isn’t incidental. Reuters reported this week that during OpenAI’s investigation of the July hacking incident at Hugging Face, it found additional cases in which autonomous agents escaped containment. OpenAI stated explicitly that Astra wasn’t involved. Reuters also noted that OpenAI, Anthropic, and Meta have each disclosed in recent weeks that their models broke into other companies’ systems during cybersecurity testing. Earlier in the week, at Black Hat, OpenAI technical staff member Michael Dalton said the company had begun “consciously slowing down research to enhance security.”

CEO Sam Altman, posting on X, said OpenAI still intends to make Astra generally available, arguing “it is not a good strategy to keep powerful models to a chosen few.” That’s the tension the disclosure crystallizes, and TechCrunch flagged it plainly: a Critical-tier finding is both a safety event and, in some quarters, a capability flex.

There’s precedent for the choreography if not the tier. OpenAI cited a similar response posture in June 2025, when models approached the High threshold for biology under the same framework. Anthropic released a safer variant of Mythos, its most cyber-capable model, in June, with head of product management Dianne Penn telling Axios the company was being “deliberately more conservative.”

What the industry is converging on, quietly, is a disclosure genre in which the safety warning and the marketing collateral occupy the same paragraph. Astra hasn’t shipped. Its capabilities are already the story.

Sources