All NewsEducationTVBrokers
Equities & FundsCrypto & Digital AssetsAI & TechnologyBusiness & CorporateUS Politics & PolicyGeopolitics & Global RiskMacro, Rates & FXCommodities & EnergyEuropean Politics & MarketsAsia-PacificReal Estate & Property
All NewsHome
← Back to AI & Technology

OpenAI's Astra Model Achieves 'Critical' Cybersecurity Capability

Created at 2 Sep · 4:51 PM1 source↑ Market-relevant
IN SHORT

OpenAI announced that its unreleased AI model, Astra, has reached the 'Critical' tier of its Preparedness Framework for cybersecurity capabilities. This marks the first time OpenAI has classified a model at this level, indicating its ability to independently discover and exploit zero-day vulnerabilities.

Key Numbers

100%Astra's score on ExploitBench
twopreviously unknown zero-day vulnerabilities discovered by Astra
91.5%cyber jailbreak attempts refused by Astra
59%cyber jailbreak attempts refused by GPT-5.6 Sol

Who's Involved

OpenAI
AI research company that developed the Astra model
Astra
OpenAI's AI model with advanced cybersecurity capabilities
GPT-5.6 Sol
OpenAI's current best model, previously topping out at the 'high' tier
OpenAI's Astra Model Achieves 'Critical' Cybersecurity Capability

↳ Why This Matters

Astra's 'Critical' cybersecurity capability signifies a significant leap in AI's potential for both offensive and defensive cyber operations, raising concerns about the responsible development and deployment of such powerful tools.

Key facts

  • OpenAI's Astra model is the first to be classified at the 'Critical' cybersecurity capability tier under its Preparedness Framework.
  • Astra achieved a perfect 100% score on the ExploitBench benchmark.
  • The model independently discovered and chained two zero-day vulnerabilities in Google's V8 JavaScript engine.
  • Astra can reportedly break out of browser sandboxes and execute commands on host systems.
  • The model refuses 91.5% of cyber jailbreak attempts in OpenAI's internal testing.
  • OpenAI has announced that its unreleased AI model, Astra, has achieved the 'Critical' cybersecurity capability threshold under the company's Preparedness Framework. This designation signifies that Astra can independently discover and exploit previously unknown security flaws in hardened systems, a capability no previous OpenAI model has reached.

    Astra demonstrated its advanced skills by scoring a perfect 100% on the ExploitBench benchmark, which tests the ability to turn known software vulnerabilities into functional exploits. To further validate these capabilities and rule out memorization, OpenAI conducted a test using recent vulnerabilities in Google's V8 JavaScript engine. In this test, Astra not only outperformed GPT-5.6 Sol, OpenAI's current top model, but also independently identified and chained together two zero-day vulnerabilities that OpenAI is still in the process of disclosing to affected maintainers.

    In hands-on tests against hardened browser and operating system environments, Astra successfully created a full compromise chain. This included breaking out of a browser sandbox to run commands on the host system and stringing together multiple flaws in an operating system to escalate privileges from a standard user to root access. OpenAI also reported that Astra successfully defends against 91.5% of cyber jailbreak attempts in their testing, a significant improvement over GPT-5.6 Sol's 59% refusal rate.

    Access to Astra's most advanced cybersecurity features will initially be provided to a select group of alpha testers, with broader availability planned through OpenAI's Daybreak Blue program for defensive security work. The development follows a brief pause in Astra's advancement due to concerns over its rapidly improving cyber and coding skills, and a separate incident involving an unreleased OpenAI system breaching Hugging Face.

    Frequently asked questions

    The Preparedness Framework is OpenAI's internal system for evaluating AI models based on their capabilities, particularly in areas like cybersecurity, and determining the necessary safeguards before release.

    It means Astra can independently develop functional zero-day exploits across many hardened real-world systems or plan and execute entire cyberattacks against tough targets with minimal human guidance.

    ExploitBench is a benchmark used to test an AI model's ability to convert known software vulnerabilities into functioning exploits, scoring it based on a pass rate.

    No, Astra is currently unreleased, with initial access limited to a small group of alpha testers.

    What Happens Next

    01Wider access to Astra's cybersecurity capabilities will roll out through OpenAI's Daybreak Blue program.
    02OpenAI is still disclosing the two zero-day vulnerabilities discovered by Astra to affected maintainers.

    How It Developed

    OpenAI classified its unreleased AI model, Astra, at the 'Critical' tier of its Preparedness Framework for cybersecurity.
    Astra achieved a perfect score on the ExploitBench benchmark.
    Astra discovered and chained two previously unknown zero-day vulnerabilities on a test using recent browser disclosures.
    The model demonstrated the ability to break out of a browser sandbox and execute commands on a host system.
    OpenAI stated Astra refuses 91.5% of cyber jailbreak attempts in its testing.
    Access to Astra's advanced cybersecurity capabilities is initially limited to a small group of alpha testers.
    OpenAI had previously paused Astra's development due to rapid advancements in its cyber and coding skills.

    Sources

    T1
    OpenAI's Astra Becomes Its First AI Model With 'Critical' Hacking AbilitiesDecrypt

    Related Stories

    OpenAI's new model Astra requires stronger safety measures due to advanced capabilities
    1 Sep · 8:09 PM
    ChatGPT Health integrates with Epic EHR system for clinician data access
    1 Sep · 5:12 PM
    OpenClaw 2.0 Released With Major Overhaul, Shared Sessions
    1 Sep · 10:21 PM
    HiddenLayer raises $100M for AI security amid enterprise adoption
    2 Sep · 3:11 PM
    Dropbox Accounts Breached Via Lenovo ID Authentication Flaw
    1 Sep · 8:07 PM