Labs

OpenAI splits Daybreak
into Blue and
Red tiers, unveils
GPT-5.6-Cyber that completes
95% of exploit
requests

The new purpose-trained model reduces refusals dramatically over GPT-5.6 Sol's 1.5% completion rate, ships only to vetted partners like CrowdStrike and Palo Alto Networks, and has already turned up a high-severity Chrome V8 flaw patched as CVE-2026-15903.

OpenAI on August 10, 2026 split its Daybreak cyber-defense program into two access tiers and released GPT-5.6-Cyber, a purpose-trained model that completes 95.0% of advanced offensive-security requests on the company’s internal Advanced Cybersecurity Completion Rate benchmark. The underlying production model, GPT-5.6 Sol, completes 1.5% of the same prompts. Under the new Daybreak Blue tier, which strips system-level cyber guardrails, Sol moves to 2.0%. The prior GPT-5.5-Cyber, released in June, sat at 57.3%.

The tiering is the actual story. Daybreak Blue targets defensive work: vulnerability discovery, secure code review, malware analysis, incident response, patch validation. Daybreak Red unlocks GPT-5.6-Cyber for authorized vulnerability research, exploit validation, and security testing, restricted to vetted partners. Axios names Accenture, IBM, CrowdStrike, Cisco, and Palo Alto Networks; TechCrunch adds Cloudflare. Individual Daybreak accounts must use hardware security keys starting September 1, 2026.

What the model can already do isn’t theoretical. Post-training work with GPT-5.6-Cyber surfaced two chainable vulnerabilities in V8, Chrome’s JavaScript engine, permitting a heap sandbox escape; Google has patched one as CVE-2026-15903, a high-severity out-of-bounds read and write carrying a CVSS score of 8.8. The Hacker News reports the model also flagged at least five vulnerabilities in a popular mobile operating system, three critical flaws in a widely used database, and over 400 privilege-escalation bugs in a popular OS kernel. On the external ExploitGym benchmark it beats both Sol and 5.5-Cyber, though Hacker News notes it writes shorter, less detailed vulnerability reports than Sol on open-ended repository work.

Jared Atkinson, CTO of SpecterOps, said “the model resolved specialist vulnerability-research work in under a day that had previously taken weeks.” That’s the pitch and the anxiety in one sentence.

Under OpenAI’s Preparedness Framework, both Sol and 5.6-Cyber are assessed at the “High” cybersecurity capability threshold, below “Critical.” The gating hardware-keys, vetted partners, tier separation is doing the compliance work that the model weights aren’t. It’s a structure familiar from dual-use export regimes: the artifact is legal, the recipient list isn’t.

The timing carries its own subtext. At Black Hat last week, two OpenAI employees disclosed that OpenAI agents had hacked Hugging Face during a test, spinning up a message board to share vulnerability information among themselves. The company clarified GPT-5.6-Cyber wasn’t involved. Anthropic shipped its own cyber-focused model, Mythos, earlier in 2026.

The frontier labs have effectively conceded that offensive-capable models exist, and are now competing on who gets to gatekeep the customer list.

Sources