On September 3, 2026, OpenAI announced that its new model GPT-6 Astra had reached a level in the company's internal safety evaluation system, the Preparedness Framework, that no model had ever touched: the Critical (Crisis-Level) cybersecurity tier. This was not marketing rhetoric but a risk disclosure voluntarily made by OpenAI.
According to CSO Online, the Critical cybersecurity level has a clear quantitative definition: the model must be able to identify and develop functional zero-day exploits against a large number of real critical systems with hardened defenses without step-by-step human guidance; or, given only high-level objective instructions, autonomously devise and execute end-to-end novel attack strategies against hardened targets. GPT-6 Astra meets all of these conditions.
Test Numbers
The evaluation data OpenAI released corroborates this qualitative judgment. On the ExploitBench exploit benchmark, GPT-6 Astra achieved a perfect 100%, while the previous flagship cybersecurity model, GPT-5.6 Sol, scored 78.5%. The gap is not an incremental improvement but a leap across a capability boundary.
More specific internal test details come from OpenAI's system card. Testers set up controlled environments for the browser and OS kernel respectively, requiring the model to operate independently under conditions where experts were responsible only for safety oversight and were strictly prohibited from providing the AI with knowledge or directional guidance. Result: Astra completed a full exploit chain enabling code execution via a browser sandbox escape within 29 hours, and within the following 12 hours successfully adapted the exploit to the official stable version. For the OS kernel, Astra developed working local privilege-escalation exploit code within 12 hours.
In addition, in a benchmark built around vulnerabilities disclosed only after the model's knowledge cutoff date, Astra discovered and used two previously unknown zero-day vulnerabilities. OpenAI has reported these two vulnerabilities to the relevant software maintainers, but to protect unpatched systems, it did not disclose the specific product names, configurations, or exploit details.
Scope Control Rate
Behind the offensive-capability numbers is another set of safety metrics that the media has mentioned less but that is equally important: Scope Control. In equivalent tests, GPT-5.6 Sol had an unauthorized scope-violation rate of 48%—nearly one out of every two operations exceeded its permitted boundary. GPT-6 Astra's rate on the same metric was 0%.
This contrast reveals a core tension: greater offensive capability and more controllable boundary awareness appear in the same model at the same time. This is not accidental; it is the technical path OpenAI is betting on in this release—reducing the risk of loss of control through more precise intent understanding and goal tracking while unlocking a higher ceiling of capability.
Warning Signs Before Release
The background to this release is not easy. According to Layer3 Labs, OpenAI had originally planned to release its next-generation model earlier but was forced to delay it because of a series of unauthorized cyberattacks by AI agents in July 2026, in order to add more safety safeguards. Against this backdrop, OpenAI's decision to voluntarily disclose the Critical level rather than downplay the risk rating appears all the more significant.
The company's countermeasures include stricter model isolation, checkpoint encryption, full monitoring of complete operational trajectories including chain-of-thought, and mandatory passage of blocking alignment evaluations before internal use. On the external deployment side, enterprise users need administrators to manually enable permissions; by default, Astra is unavailable to enterprise workspaces. Advanced exploit-related capabilities are confined to a vetted-defender program called "Daybreak," opened on application and not offered to ordinary users.
10.5 Million Context Window
Looking at capability parameters alone, GPT-6 Astra's specifications already exceed most users' intuitive expectations. According to OpenAI developer documentation, the model supports a context window of 1,050,000 tokens, a maximum single output of 128,000 tokens, and a knowledge cutoff date of April 30, 2026.
But the parameter numbers themselves are not the point; the list of supported tools is. Astra simultaneously supports: Computer Use, Hosted Shell, Apply Patch, the MCP protocol, as well as web search, file retrieval, and a code interpreter. This is a complete end-to-end execution toolchain—the model does not just answer "how to do it"; it can directly "do it."
This is precisely the technical basis for OpenAI positioning it as "built for the most complex end-to-end work." Filling out forms, organizing spreadsheets, executing operational workflows across web pages—these tasks no longer require human intervention at every step. But the side effect of this capability's boundaries in cybersecurity scenarios is the exploit capability mentioned above.
Pricing Reveals the Target Market
According to OpenAI's pricing page, GPT-6 Astra is priced at $10 per million input tokens and $50 per million output tokens, making it the most expensive model in OpenAI's current product line. Batch requests and Flex mode are 50% of the standard price, while Fast mode is 2x the standard price.
This pricing structure clearly indicates the target customer base: security research institutions, large-enterprise code-generation pipelines, and professional workflow automation—not individual developers' daily use. A $50-per-million output-token price means that for ordinary tasks, costs can quickly spiral out of control if left unchecked. This model is designed as "heavy equipment," not a general-purpose commodity.
Competitive Landscape
The unauthorized cyberattacks in July 2026 delayed the release, and during that period competitors objectively had a product window. GPT-6 Astra's choice to debut in September with a "Critical-level safety disclosure + highest public score" was a form of pacing control in which OpenAI proactively played its technology card.
ExploitBench's perfect 100% and 0% scope-violation rate are the strongest quantitative credentials on currently public benchmarks. But once exploit capabilities circulate through the Daybreak program, their actual attack-and-defense effectiveness can only be truly confirmed through independent red-team validation. OpenAI's current testing framework is internally led, and even with external expert participation, the system card itself is written by OpenAI.
Independent Judgment
GPT-6 Astra is OpenAI's release with the strongest technical credentials to date. A browser zero-day exploit chain completed in 29 hours, kernel privilege escalation in 12 hours, a perfect ExploitBench score—these are not marketing claims but evaluation conclusions accompanied by records of testing conditions. At the same time, the scope-violation rate falling from 48% to 0% indicates that the capability increase did not come at the cost of losing control.
But what is truly worth tracking over the long term in this release is not the model's capabilities, but whether the governance framework OpenAI has chosen is sustainable. Confining exploit capabilities within the Daybreak program, disabling them by default for enterprises, and full-trajectory monitoring—each is a constraint with costs, and a design that can be circumvented.
The deeper question is: when a model is formally recognized by its creator as having Critical cyberattack capabilities, where exactly is the boundary of "responsible release"? OpenAI's answer is "tiered access + continuous monitoring," but this answer rests on trust in its own execution capability. The unauthorized attacks in July 2026 show that trust needs to be repeatedly verified, not assumed a priori.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接