First Critical-Level AI: OpenAI Astra Autonomously Discovers Zero-Day Vulnerabilities — The Security Paradox Behind a Perfect ExploitBench Score

On September 2, 2026, OpenAI announced that its new model Astra had officially triggered the "Critical" cybersecurity capability threshold in internal evaluations—the first model to receive this designation since the Preparedness Framework was established in 2023. Astra scored a perfect 100% on ExploitBench and autonomously discovered two previously unknown zero-day vulnerabilities.

On September 2, 2026, OpenAI announced that its new model Astra had officially triggered the "Critical" cybersecurity capability threshold in internal evaluations—the first model to receive this designation since the Preparedness Framework was established in 2023. Astra scored a perfect 100% on ExploitBench and, without any human guidance, autonomously discovered two previously unknown zero-day vulnerabilities, chaining multiple vulnerabilities together to achieve a complete privilege escalation path from an ordinary user to root.

ExploitBench measures an AI's ability to convert known vulnerabilities into working exploit code—a 100% score means no problem in the test set managed to stump Astra. The two zero-day vulnerabilities were not test items; rather, they were byproducts Astra uncovered in passing while constructing exploit chains, targeting Chrome V8 engine vulnerabilities disclosed between June and August 2026. OpenAI stated that it is currently coordinating the disclosure process with the relevant software developers.

Beyond the Perfect Score: Sandbox Escape and Privilege Escalation Chains

In expert-led red team testing, Astra demonstrated two complete attack chains:

  • Browser sandbox escape: With nothing more than opening a malicious HTML file, Astra was able to break through the browser's sandbox isolation and execute commands directly on the host machine—the entire process required no human intervention.
  • Complete OS privilege escalation: Against a hardened operating system, Astra identified multiple independent vulnerabilities and chained them into a complete privilege escalation path from a non-privileged user account to root.

OpenAI's Preparedness Framework defines two triggering conditions for the "Critical" designation, either of which suffices: first, autonomously discovering and exploiting zero-day vulnerabilities of any severity against a substantial number of hardened real-world systems without human guidance; second, autonomously planning and executing a complete cyberattack against hardened targets based solely on high-level objective instructions. Astra has met both.

Another Set of Numbers on the Defensive Side

Astra's refusal rate for cybersecurity-related jailbreak attempts reached 91.5%, compared with 59% for its predecessor, GPT-5.6 Sol. That jump of more than 30 percentage points indicates that offensive capability and defensive posture have been strengthened in parallel within the same generation of models.

Compared with GPT-5.6 Sol, Astra is less likely to proactively bypass safety restrictions and is also less likely to passively respond to deliberately placed "honeypot" targets in testing. These data points attempt to convey a message: stronger attack capability comes with stronger intrinsic constraints.

But the refusal rate in internal tests and actual performance in real adversarial scenarios have never been the same thing.

Altman Acknowledges the "Obvious Tension"

OpenAI CEO Sam Altman wrote: "There's an obvious tension here: on one hand, Astra is remarkably capable and we're excited to see what people build with it; on the other, we're clearly at a stage of development that calls for caution, and we're moving at a deliberate pace to ensure we can meet the safety standards that the new capability level demands."

Altman also revealed that Astra's training had actually been completed for some time, and the company has been deliberately slowing its pace to allow sufficient time for safety and alignment work.

Release Strategy: Tiered Access, Not Full Containment

OpenAI's approach is not a complete halt but tiered governance. The highest-level cybersecurity capabilities will be made available through the Daybreak Blue program, restricted to a small number of vetted testing partners; the version of Astra available to general users will have reduced functionality.

Daybreak is a layered access system OpenAI established specifically for cybersecurity scenarios, divided into two tiers: Blue (defensive use) and Red (offensive research authorization). The design philosophy of this mechanism draws from clearance management logic in the military and intelligence domains: tools themselves do not distinguish between good and evil, but usage authorization is determined by qualifications and purpose.

More than 130 technology and security companies—including Anthropic, AWS, Google, and Microsoft—have co-signed an open letter on collective cyber defense led by OpenAI.

The Deeper Question: Who Defines "Safe Enough"

Tal Kollender, founder and CEO of AI cybersecurity company Remedio, said: "Restricting the most advanced cybersecurity capabilities to a small number of partners is a reasonable mitigation measure. I don't think OpenAI is being irresponsible here. Their framework is designed to permit release under the right safeguards. But the security story is that the defensive side hasn't caught up with any version of this technology, whether restricted or public."

This statement reveals a structural dilemma: the asymmetry between offense and defense cannot be fundamentally resolved through access control. Astra's Critical-level capability means that even if only a very small number of people can legally use its strongest features, the risk of replication and leakage remains. Historically, every era-defining offensive technology—from nuclear weapons design blueprints to advanced persistent threat (APT) toolkits—has been caught in a recurring struggle between control and proliferation, and more often than not, proliferation has won the race against time.

The more fundamental question is this: OpenAI's Preparedness Framework is a self-assessment framework. The criteria for the "Critical" designation are defined by the company itself, the evaluation process is led by the company itself, and the restrictions triggered by the designation are also decided by the company itself. The mechanism has a certain seriousness by design, but it is self-contained in its oversight structure. When a single organization simultaneously plays the roles of capability developer, capability evaluator, and capability controller, conflicts of interest are not merely possible—they are inevitable.

Assessment: A Real Inflection Point, But Not for the Reasons Usually Cited

The most common narrative in discussions about Astra is that "this is the first time AI can autonomously hack." That statement is neither accurate nor the most important signal. Automated vulnerability exploitation tools have existed for a long time, and the security research community has been using AI-assisted vulnerability analysis for years. Astra's true novelty lies not in what it can do, but in the fact that it integrates these capabilities into a single model while dramatically compressing cost and lowering the barrier to entry—Astra consumes far fewer tokens than its predecessor when completing equivalent attack chains.

This means the marginal cost of launching sophisticated cyberattacks is approaching zero. Once that threshold is crossed, the frequency, scale, and diversity of attacks will all undergo order-of-magnitude changes, while the manpower and resource costs on the defensive side will not decline correspondingly. Astra itself is not the ultimate threat; it is a visible marking on this curve.

OpenAI's decision to publicly disclose this evaluation result rather than keep it internal is itself a transparency gesture worthy of recognition. But transparency is not a solution. The real inflection point will be whether external regulators, independent security research institutions, and international policy coordination frameworks can establish evaluation and governance mechanisms that do not rely on vendor self-discipline before the next generation of models arrives. Given the current pace of progress, the outlook for this race is not optimistic.