Key facts
- OpenAI's Astra model is the first to be classified at the 'Critical' cybersecurity capability tier under its Preparedness Framework.
OpenAI announced that its unreleased AI model, Astra, has reached the 'Critical' tier of its Preparedness Framework for cybersecurity capabilities. This marks the first time OpenAI has classified a model at this level, indicating its ability to independently discover and exploit zero-day vulnerabilities.

Astra's 'Critical' cybersecurity capability signifies a significant leap in AI's potential for both offensive and defensive cyber operations, raising concerns about the responsible development and deployment of such powerful tools.
OpenAI has announced that its unreleased AI model, Astra, has achieved the 'Critical' cybersecurity capability threshold under the company's Preparedness Framework. This designation signifies that Astra can independently discover and exploit previously unknown security flaws in hardened systems, a capability no previous OpenAI model has reached.
Astra demonstrated its advanced skills by scoring a perfect 100% on the ExploitBench benchmark, which tests the ability to turn known software vulnerabilities into functional exploits. To further validate these capabilities and rule out memorization, OpenAI conducted a test using recent vulnerabilities in Google's V8 JavaScript engine. In this test, Astra not only outperformed GPT-5.6 Sol, OpenAI's current top model, but also independently identified and chained together two zero-day vulnerabilities that OpenAI is still in the process of disclosing to affected maintainers.
In hands-on tests against hardened browser and operating system environments, Astra successfully created a full compromise chain. This included breaking out of a browser sandbox to run commands on the host system and stringing together multiple flaws in an operating system to escalate privileges from a standard user to root access. OpenAI also reported that Astra successfully defends against 91.5% of cyber jailbreak attempts in their testing, a significant improvement over GPT-5.6 Sol's 59% refusal rate.
Access to Astra's most advanced cybersecurity features will initially be provided to a select group of alpha testers, with broader availability planned through OpenAI's Daybreak Blue program for defensive security work. The development follows a brief pause in Astra's advancement due to concerns over its rapidly improving cyber and coding skills, and a separate incident involving an unreleased OpenAI system breaching Hugging Face.