OpenAI Astra Reaches ‘Critical’ Cybersecurity Level

OpenAI’s Astra was classified at the Preparedness Framework’s ‘Critical’ level after internal tests showed it can autonomously find and exploit zero-day vulnerabilities.

OpenAI announced that its newest model, Astra, has reached the Preparedness Framework’s ‘Critical’ cybersecurity capability level, the first of the company’s models to receive that designation. The classification applies when a model can independently find and exploit zero-day vulnerabilities across many well-defended systems or carry out a complete cyberattack against a hardened target from a high-level instruction.

In internal evaluations, Astra scored a perfect result on ExploitBench, a benchmark that measures a model’s ability to turn known vulnerabilities into working exploits. In a separate test focused on more recently disclosed flaws, the model independently discovered two zero-day vulnerabilities. OpenAI reported that Astra escaped a browser sandbox to run commands on the underlying machine and chained multiple flaws in a hardened operating system to obtain root access.

On safety metrics, Astra declined 91.5% of cyber-related jailbreak attempts in the company’s testing, compared with a 59% decline rate for its predecessor, GPT-5.6 Sol. OpenAI reported that Astra showed a lower tendency than Sol to bypass safety restrictions or to exploit deliberately placed honeypot targets during evaluations.

The company said the ‘Critical’ classification requires extra safeguards before broad release. OpenAI plans to give a group of testers early access to Astra’s cybersecurity capabilities and to expand availability through its Daybreak Blue program. Full cybersecurity features will not be widely available at launch.

OpenAI wrote in a statement: “We are entering a stage of AI development in which models can take on more consequential work, and failures of alignment and control can have more serious effects. Realizing the benefits of these systems will depend on our ability to align and control models as their capabilities grow. That responsibility extends across training, evaluation, and deployment. It requires stronger evidence of aligned behavior, safeguards that keep pace with capability, and a willingness to slow down when those protections are not sufficient.”

Nearly 130 technology and cybersecurity companies have expressed support for an OpenAI-led initiative to strengthen cyber defenses in response to increasingly sophisticated AI-enabled attacks. OpenAI has updated its security measures, including sandboxing, faster alerting, and pausing training when necessary, to manage risks as models develop stronger offensive capabilities.

Under the Preparedness Framework, the ‘Critical’ level triggers heightened review and control measures. OpenAI said the designation followed controlled tests that demonstrated the model’s ability to autonomously exploit vulnerabilities and perform advanced attack techniques, prompting the company to require extra protections and a staged rollout.

Articles by this author