- Lead. OpenAI disclosed on September 1 that its forthcoming Astra model is the first system it has ever rated “Critical” under the company’s Preparedness Framework—a threshold that, by the framework’s own terms, requires additional safeguards before any release.
- Fact. In internal testing, Astra scored 100% on ExploitBench, a benchmark for exploit development from known vulnerabilities, and independently discovered and exploited two zero-day vulnerabilities from a set of twenty high-severity issues disclosed between June and August 2026.
- Stake. The disclosure raises a pointed question about whether AI safety frameworks designed to pace development are doing so, or functioning primarily as disclosure mechanisms after the fact.
Under OpenAI’s Preparedness Framework, a model reaches the Critical cybersecurity tier if it can “identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention,” or devise and execute novel cyberattack strategies against hardened targets given only a high-level goal. Astra, disclosed in a September 1 blog post, met both criteria in internal evaluation.
What the Tests Showed
On ExploitBench, Astra achieved a perfect score, demonstrating its ability to turn known vulnerability disclosures into working attack code. On a separate internal benchmark involving twenty high-severity vulnerabilities disclosed between June and August 2026, Astra did not merely exploit existing proof-of-concept code—it incorporated two previously undisclosed zero-days into an exploit chain on its own, without human direction at intermediate steps.
OpenAI described the performance as “the first time” a model in its portfolio has crossed the Critical threshold. The Preparedness Framework was introduced in 2023 as a self-regulatory commitment to evaluate models across four risk categories—cybersecurity, biological, nuclear and radiological, and persuasion—before release. The cybersecurity rubric defines Critical as the ceiling tier. Models rated below Critical, at “High,” are permitted for release with enhanced monitoring; the Critical rating formally requires a higher bar of demonstrated safeguard effectiveness.
The Release Decision
OpenAI said it believes Astra’s safeguards “sufficiently minimize the risk of severe harm for release” but acknowledged that access to its cybersecurity capabilities will be more limited than for its standard model tiers. The company did not give a specific launch date. Access to the features underpinning the ExploitBench performance—autonomous vulnerability discovery and chaining—will be restricted to vetted enterprise customers, government partners, and security researchers, according to reporting by Axios.
The decision to release a Critical-rated model, even with restrictions, is notable. It suggests that OpenAI considers the competitive and commercial pressure to release frontier models to be manageable within a restricted-access architecture, rather than grounds for an indefinite hold. The pace of OpenAI’s recent product decisions—including the August 30 removal of DALL-E and the acceleration of GPT Image 1—points to a company moving faster on its product roadmap even as it navigates its own safety assessments.
Reactions and Oversight Questions
Security researchers responded with a mix of concern and cautious acknowledgement that OpenAI had, at minimum, disclosed the capability publicly before release. The company’s previous frontier models had reached the High tier in cybersecurity, so the Critical designation represents a material step in capability. Absent an independent audit of OpenAI’s internal benchmarks, the security community cannot verify whether the ExploitBench scores translate to real-world attack capability at the level the framework implies.
Congressional staff familiar with AI oversight policy noted that no current US regulation requires companies to disclose Preparedness Framework scores or pause releases at the Critical tier. The disclosure functions as a voluntary commitment with no external enforcement mechanism—a gap that advocates for statutory AI governance have cited as the central weakness of the industry’s current self-regulatory posture.