OpenAI has revealed that its forthcoming Astra model meets the “Critical” cybersecurity capability threshold under its Preparedness Framework, marking a significant milestone in AI development that raises both defensive possibilities and serious security concerns.
What “Critical” Means
The “Critical” designation applies when an AI model can independently detect and exploit zero-day vulnerabilities across multiple well-defended systems, or execute a complete cyber attack against a hardened target from only a high-level instruction without human guidance along the way. Astra has demonstrated exactly these capabilities during controlled evaluations.
In testing, Astra achieved a perfect 100% score on ExploitBench, a benchmark for developing exploits from known vulnerabilities. The model discovered and used two previously unknown vulnerabilities in unspecified software as part of an exploit chain, and found multiple vulnerabilities in a hardened operating system, combining them into a local privilege-escalation chain from an unprivileged user to root.
Enhanced Safeguards
The disclosure triggered additional controls around access, monitoring, and deployment. OpenAI says it has added classifiers and layered protections to improve robustness against misuse and prevent unauthorized model actions, even in the absence of a malicious user.
“We believe Astra’s safeguards sufficiently minimize the risk of severe harm for release under our Preparedness Framework,” the company stated. However, OpenAI acknowledged that Astra’s enhanced guardrails may also erroneously flag legitimate security research as misuse—a challenge for defenders who need access to these capabilities.
Astra will initially be restricted to testers through the Daybreak Blue program, similar to how Google and Anthropic have restricted their most capable cyber models.
The Broader Context
The development comes amid a wave of frontier AI cybersecurity model releases. Google unveiled Gemini 3.8 Flash Cyber, claiming it surpasses larger frontier models in autonomous vulnerability discovery. Anthropic released Claude Mythos 5.1 with enhanced safeguards. Over 100 companies, including Anthropic, Google, Microsoft, and OpenAI, have signed a joint letter calling for improved defenses against AI-fueled cyber attacks.
The practical implications are significant: the same systems that help defenders find weaknesses faster could dramatically lower the expertise and labor needed to conduct sophisticated attacks. How labs balance these competing uses will define the next chapter of AI security.