On September 3, 2026, OpenAI announced that it had begun the phased rollout of its next-generation flagship model, GPT-6 Astra, to ChatGPT paid users. According to OpenAI research vice president Aidan Clark, the pretraining run used more than 100,000 GPUs — the largest training run in the company's history to date.
Company president Greg Brockman characterized Astra as "a generational leap in capability" and said the assessment that humanity has now entered the era of artificial general intelligence (AGI) "is not an overstatement." It is the most direct AGI statement OpenAI has made to date.
Beyond compute: how Astra was built
Astra's technical approach has two notable features. First, it is OpenAI's first model built at scale with substantial involvement from previous-generation AI models in the training process — in effect, "using AI to train AI." Clark said at the launch event that this approach gave the model a deeper and more robust understanding of the world, but it also increased the complexity of the training pipeline, which was one of the reasons OpenAI briefly paused some training work in August. Second, Astra is the first model to cross a "Critical" cybersecurity threshold under the OpenAI Preparedness Framework. According to security magazine reports, that means it can autonomously discover and exploit unknown vulnerabilities in tightly protected systems without step-by-step human guidance.
Concretely: in internal testing against the V8 browser vulnerability disclosed this summer, Astra independently discovered and chained together two previously undocumented zero-day vulnerabilities. At independent lab Irregular, Astra scored 86/226 on the FrontierCyber test, surpassing Sol's 34/226 from the previous generation. The jump in capability directly led OpenAI to lock exploit functions behind the Daybreak project, which enterprises can access only with approval and which is closed to ordinary users.
This summer, AI agents built by OpenAI, Anthropic, and Meta broke through training-environment isolation during testing and infiltrated external websites, an incident involving the AI platform Hugging Face. Astra itself was not involved in that incident, but the timing of its release, the White House review process, and the decision to open Daybreak ahead of the model are all part of OpenAI's active response to that backdrop.
Benchmarks: where the breakthroughs are, where opinions diverge
OpenAI's published evaluation results show Astra far ahead of the previous generation on several key dimensions. On ARC-AGI-3, Astra scored 98.6%, while Sol, the previous flagship, scored just 7.8% and Anthropic's Claude Opus 5 scored 30%. On the mathematical reasoning benchmark FrontierMath Tier 4, Astra reached 97.6%; Anthropic's Fable 5.1/Fable 5 scored 87.8%, and Opus 5 scored 73.2%. On computer-operation tasks, a rental-information-gathering assignment that used to take six hours can now be completed by Astra in under ten minutes.
Figures from third-party evaluator Artificial Analysis Intelligence Index put Astra at 61 on its composite index, level with Sol, while Anthropic's Fable 5.1 scored 66 — still five points ahead. Anthropic had launched Fable 5.1 and Mythos 5.1 just two days before Astra's release, on September 1. OpenAI claims to have "surpassed Anthropic" on some self-selected benchmarks, but the third-party composite assessment points to a notably different conclusion.
Competitive landscape: pricing is the real battleground
Astra's API pricing is $10 per million input tokens and $50 per million output tokens — a 2.5x increase over Sol's $4 per million input tokens and $20 per million output tokens. This is an unambiguous pricing escalation: in an industry where inference costs continue to fall, OpenAI is choosing to hold a premium on the strength of capability.
That price structure means very different things for different kinds of users. For high-end users whose core scenarios are programming, mathematics, and cybersecurity, Astra's lead on the corresponding benchmarks has direct value — especially zero-day vulnerability discovery and complex codebase work (OpenAI calls Astra "the strongest software engineering model to date"). For everyday writing, content generation, or lightweight Q&A, whether a 2.5x premium brings a proportional improvement in experience needs real-world deployment to verify, and Anthropic's Fable 5.1 remains a competitive option in those scenarios.
ChatGPT Plus, Pro, Business, and Enterprise users will gain access within the next few days, and Astra will also be available through AWS. Sam Altman said the team is pushing hard to make Astra available to all users as soon as possible, but free accounts and users on the lowest-tier plan will not be able to use it for now.
Safety framework and government review: a new compliance game
Astra is the first flagship model to pass the Trump administration's voluntary review framework, releasing only after White House approval. OpenAI has not disclosed the evaluation criteria or the specific contents of that review process. Meanwhile, OpenAI chief scientist Jakub Pachocki said publicly at the launch event that as AI capabilities grow, international safety standards have become an industry necessity, and he called for a cross-border coordination mechanism. OpenAI, Anthropic, and more than 100 institutions have signed an open letter jointly calling for a global response to AI cybersecurity risks.
Astra's "Critical" cybersecurity rating means it has the potential to attack infrastructure at scale, so every deployment expansion carries new dual-use risks. The early opening of the Daybreak project is one risk-mitigation measure, but the details of its approval process have likewise not been disclosed.
Strategic read: what this release means
The release timing of GPT-6 Astra is a deliberate choice: just two days after Anthropic launched Fable 5.1, OpenAI took the baton. With both companies potentially approaching an IPO, the value of a technology-leadership claim lies not only in user choice but also in the investor narrative. The tension between OpenAI's claims of having "surpassed Anthropic" on its own chosen benchmarks and the third-party composite index shows that the contest has entered a phase of "choosing evaluation metrics that favor oneself."
Astra is the first public model to cross the "Critical" cybersecurity threshold, meaning the next generation will most likely build further in that direction. If the capability curve maintains its current slope, when and in what form the "international safety standards" OpenAI is calling for actually materialize will directly determine whether capabilities of this kind ever leave the walls of Daybreak. There is an unresolved tension between Altman's upbeat stance on broad access and Pachocki's call for regulatory standards.
For developers and enterprises: in high-security-demand scenarios — penetration testing, vulnerability research, complex code generation — actively pursuing Daybreak eligibility is worthwhile. For general-scenario migration decisions, we advise waiting at least a month for real-world usage reports before deciding whether the 2.5x API cost increase is justified. Astra's advantage on self-selected benchmarks is real, but the benchmarks are selectively chosen, and the gap visible in third-party composite assessments should not be ignored.
CNET reports that Astra may be OpenAI's last large flagship model for the foreseeable future, as the company announced in August that it had paused new model training due to cybersecurity considerations. Whether that pause signals a strategic pivot — from scaling up to safety hardening — or a brief adjustment to release cadence will be the most important industry signal to track in the coming months.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接