OpenAI has a document called the Preparedness Framework that sorts a model’s most dangerous capabilities into tiers, and until this week nothing the company had built had ever reached the top of the cybersecurity tier. On September 1, 2026, OpenAI said that had changed. Its upcoming model, known internally as Astra, is the first system the company has ever classified as meeting the Critical cybersecurity capability threshold, meaning it can find previously unknown flaws in hardened, real world systems and figure out how to exploit them without a person walking it through each step.
The proof was not a hypothetical. During evaluation, Astra scored a perfect result on ExploitBench, a benchmark built from known, already patched vulnerabilities. Then, in a modified version of that same test, it went further than the benchmark was designed to measure: it discovered two zero day vulnerabilities on its own and chained them into a working exploit. OpenAI is now in the process of disclosing both flaws to the software maintainers responsible for fixing them.
The short version
- Astra is the first OpenAI model rated Critical for cybersecurity under the company’s own Preparedness Framework
- It scored 100% on ExploitBench and found two real zero day vulnerabilities unprompted, during a modified version of that test
- OpenAI delayed the release starting August 7, adding mandatory monitoring for every inference run where Astra had access to tools
- General access is coming “soon,” but Astra’s strongest offensive cyber capabilities will not be part of it
- Advanced access is restricted to Daybreak, OpenAI’s vetted cybersecurity coalition, split into Daybreak Blue and Daybreak Red tiers
- Sam Altman’s framing: “we do not think it is a good strategy to keep powerful models to a chosen few,” but “given its cyber capabilities, we need a little bit longer to do this safely”
What “Critical” actually means inside OpenAI’s own rulebook
OpenAI’s Preparedness Framework was not written for this specific moment, but it was written to anticipate it. Under the framework’s definitions, a model crosses the Critical threshold for cybersecurity if it can identify and build functional zero day exploits across the full range of severity in many hardened, real world systems without a human directing each step, or if it can independently design and carry out an entire novel attack strategy against a hardened target starting from nothing more than a high level goal. Those are deliberately narrow, hard to hit bars. Reaching either one is supposed to be rare enough that it forces a company to stop and change how it ships the model, not just note the result in a research paper.
That is what happened here. OpenAI says it began slowing parts of Astra’s development back on August 7, once internal testing suggested the model might be approaching this tier, and added a mandatory monitoring layer over every inference call where Astra had access to external tools. The company has publicly framed the months since then as time spent strengthening safeguards against misuse and unauthorized autonomous action, rather than time spent making the model more capable. Whether that framing survives outside scrutiny is a separate question, but the sequencing itself, capability first, safeguards after, then a delayed and restricted release, mirrors a similar pattern the company documented with an earlier model that solved ten decades old math problems for around $2,000 and got flagged as a cyber risk in the same breath.
Why a perfect benchmark score is the less interesting part
A model acing ExploitBench is impressive but expected. It is a known quantity, a fixed set of already documented vulnerabilities that a well trained system can learn to exploit reliably, the AI equivalent of acing a practice exam once you have seen the answer key. What actually pushed OpenAI to make a Critical classification was what happened when testers modified that benchmark to include conditions the model had not been trained against. Astra did not just fail gracefully or refuse. It found two vulnerabilities nobody had catalogued, on real, hardened targets, and worked out how to chain them into something that functioned as an exploit, without a human providing the missing steps.
| Milestone | Detail |
|---|---|
| ExploitBench score | 100%, on known, previously catalogued exploits |
| Zero days found unprompted | Two, chained into a working exploit during a modified test |
| Human guidance required | None, for the exploit chain in question |
| Disclosure status | In progress, to the affected software maintainers |
| Preparedness Framework tier | Critical, cybersecurity category, the first OpenAI model to reach it |
| Monitoring added | Mandatory oversight on all tool-using inference, starting Aug 7, 2026 |
That distinction, known bugs versus previously undiscovered ones, is exactly the line OpenAI’s own framework was built to watch for. A model that is good at known exploits is a faster penetration tester. A model that reliably finds unknown ones in hardened systems on its own is something closer to an automated vulnerability researcher that never sleeps and does not need a security clearance to keep working, which is precisely why access to that specific capability is being kept narrow.
Daybreak, and who actually gets the powerful version
OpenAI is not withholding Astra outright. The company says a general version will be available “soon,” giving ordinary users and businesses the model’s everyday capabilities. What will not ship broadly is the offensive cyber tooling that triggered the Critical rating in the first place. That stays inside Daybreak, OpenAI’s cybersecurity coalition, which the company expanded in August into two tiers: Daybreak Blue and the more sensitive Daybreak Red, aimed at vetted defenders, researchers and partner organizations rather than the general public.
Sam Altman has tried to hold two positions at once in public comments about the release. He has said OpenAI does not think keeping powerful models locked to a small, chosen group is good strategy long term, a defense of broad access as a principle. But on Astra specifically, his own words concede the opposite in practice: “given its cyber capabilities, we need a little bit longer to do this safely.” That gap, between the stated philosophy and the actual rollout, is where most of the informed skepticism about this launch is landing. It is a similar tension to the one already playing out around Nvidia’s AI security alliance, which has tripled to 120 members without OpenAI or Anthropic actually joining it, each lab insisting it takes AI security seriously while keeping its own frontier safety work mostly in house.
The stakes if this capability gets loose
The reason a classification like this matters beyond OpenAI’s internal paperwork is the gap between a controlled lab environment and the real world, where the damage from an automated vulnerability finder is not hypothetical. Software supply chain attacks are already a live threat without AI involved: a single poisoned security scanner recently leaked 153GB of credentials from nearly 2,500 companies, damage caused by a comparatively simple, human built attack. A model that can autonomously discover and chain unknown exploits against hardened targets represents a meaningful jump in what a single actor, criminal, state sponsored or otherwise, could do with the right access, which is exactly why OpenAI is treating distribution as the actual safeguard here rather than the model’s underlying capability itself.
What comes next
OpenAI has not given a firm release date for the general version of Astra, only “soon,” and has not detailed exactly how organizations qualify for Daybreak access beyond describing it as vetted. Two things are worth watching once it does ship. First, whether the two zero day vulnerabilities Astra found get patched before any details leak publicly, since a delay in disclosure timing is exactly the kind of gap that turns a responsible research finding into a real world incident. Second, whether other frontier labs follow OpenAI in publishing their own Critical level classifications, or whether this becomes another case where one company sets a safety precedent that the rest of the industry quietly declines to match. Astra is the first model to cross this specific line. Given how fast frontier capabilities have been moving all year, it is very unlikely to be the last.

