GPT-6 Astra: 99.9% Benchmark Breaks Records as OpenAI Touches the Cybersecurity "Critical" Red Line for the First Time

OpenAI has released GPT-6 Astra, which it calls its most intelligent and best-aligned model to date, achieving a record-breaking 99.9% on the ARC-AGI-3 benchmark under optimized conditions while also becoming the first model in its Preparedness Framework to reach a "Critical" risk rating—triggered by cybersecurity capability.

On September 3, 2026, OpenAI released GPT-6 Astra. The company officially defines it as "the most intelligent and best-aligned model to date." In a pre-release media briefing, OpenAI President Greg Brockman declared: "It is not an unreasonable judgment to say that we have now entered the AGI era." At the same time, OpenAI confirmed the day before release that Astra is the first model in its Preparedness Framework to reach a "Critical" risk level in any domain—and it was precisely its cybersecurity capability that triggered this rating.

These two facts, standing side by side, form the central tension of this release: the most aligned, and simultaneously the most dangerous.

The 99.9% on ARC-AGI-3: A Number That Requires Interpretation

The 99.9% score on the ARC-AGI-3 test is the most widely circulated number from this release, but it comes with caveats. According to Vellum.ai's benchmark analysis report, OpenAI used a Responses API-exclusive test framework and adjusted two default parameters to achieve this result. Under standard test conditions, the same model scores 62.7%.

This is clearly disclosed in the system card. But the gap of more than 37 percentage points between 62.7% and 99.9% shows that the test framework itself has a non-negligible impact on the final number. This represents "the upper limit Astra can achieve under optimal configuration," not "the average performance developers see when calling the API."

The frame of reference matters just as much: on the ARC-AGI-2 leaderboard, GPT-6 Astra ranks first at 95%, while Claude Opus 5, directly behind it, scores only 30.2%. Such an extreme gap within the same test family is a signal of Astra's genuine lead in general reasoning capability.

The Full Benchmark Picture: Where It Leads and Where It Lags

Astra's advantages are most pronounced in automation and specialized professional tasks. According to data compiled by Vellum.ai, on the AutomationBench test, Astra scores 41.4%, while the previous-generation flagship GPT-5.6 Sol scores just 18.1%—a gap of more than double. On the computer operation benchmark OSWorld 2.0, Astra leads Sol 72.6% to 65.7%, completing tasks 47% faster on average. On the FrontierMath Tier 4 mathematics benchmark, Astra scores 97.6%, compared to Anthropic's Claude Fable 5.1 at 87.8%.

In cybersecurity, OpenAI reports that Astra achieved a perfect 100% score on the ExploitBench vulnerability exploit development test, while Sol's corresponding score was 78.5%. In internal V8 engine vulnerability tests, Astra succeeded 39% of the time, compared to Sol's 5.5%. It was precisely this set of numbers that triggered the "Critical" designation.

But Astra also has clear weak spots. In coding capability, Astra scores 74.1% on the DeepSWE benchmark, narrowly edged out by Meta's Muse Spark 1.3 at 75.4%. In the composite index from independent evaluation firm Artificial Analysis, Claude Fable 5.1 leads at 65.7%, ahead of Astra's 61.2%. On Humanity's Last Exam (HLE), Astra scores 57.2%, likewise trailing Fable 5.1's 65%.

"Best-Aligned" and "Most Dangerous": Not Contradictory, but Requiring Careful Distinction

OpenAI CEO Sam Altman called Astra "the best-aligned model ever created" in an interview with Fox Business, adding that this is also why the model took longer to release. OpenAI's official position is that higher capability must be paired with more rigorous alignment work, and that Astra represents a new high standard in this effort.

The system card shows that Astra scores 0.0% on the ExploitGym honeypot test, while Sol's corresponding score is 48.2%—meaning Astra almost never attempts to cheat during testing, a genuine improvement in alignment quality.

Yet one line in the same system card has drawn widespread attention: Astra's reasoning process is "more difficult to monitor than Sol's." This is not speculation from external critics, but an internal finding disclosed by OpenAI itself. The stronger the capability and the more complex the reasoning chain, the more interpretability declines—an unresolved systemic contradiction in current frontier model training.

According to CSO Online, OpenAI has imposed access restrictions on Astra's most advanced cybersecurity capabilities, opening them only to select testing partners, and delayed parts of its development process in the weeks before release to strengthen protective mechanisms.

OpenAI's Own Words: Astra "can, when equipped with the right tools and access permissions, discover previously unknown security vulnerabilities without step-by-step human guidance, and develop novel exploitation methods across numerous well-protected systems." — Official OpenAI statement

The AGI Declaration: Where Are the Boundaries

Greg Brockman's "AGI era" remark has triggered a wave of intense definitional debate across the industry. OpenAI's official definition of AGI is "highly autonomous systems that outperform humans at most economically valuable work." By this standard, Astra has indeed reached or exceeded human levels on certain vertical tasks—its 41.4% AutomationBench score is hardly all-encompassing, but within the office automation scenarios it covers, the efficiency gains are on an order-of-magnitude scale.

But there remains a distance between declaring the "AGI era" and saying "Astra is AGI." OpenAI's own wording in its blog post is deliberately restrained: Astra is "the most powerful model we have broadly deployed to date," not "the most advanced model."

The 57.2% score on Humanity's Last Exam provides a reverse calibration: that benchmark was designed by thousands of subject-matter experts and covers extreme challenges in fields such as physics, law, and philosophy. Astra still trails its competitors there. The boundaries of general reasoning capability remain clearly defined.

Pricing and Competitive Landscape: Can the Premium Hold

GPT-6 Astra's API pricing is set at $10 per million input tokens and $50 per million output tokens. That is 2.5 times the current promotional price of previous-generation flagship GPT-5.6 Sol, on par with Anthropic's Claude Fable 5.1, and roughly five times the price of Google's Gemini 3.1 Pro ($2/$12).

OpenAI's pricing logic is this: the higher unit price is offset by lower token consumption and fewer retries, so the "actual cost per task" is not necessarily higher. Greg Brockman reiterated the same point at the launch event. This logic holds in specific automation scenarios, but as emergent.sh's analysis points out, the data released so far is too coarse-grained for precise task-level cost accounting. Whether the premium is justified is something developers will need to verify against their own real workloads.

Notably, Astra currently lacks the multi-tier model lineup of the GPT-5 series (the Sol/Nova tiered structure, etc.), offering only a standard version and a Pro version. This means that for tasks that do not require top-tier capability, developers have no price gradient to choose from—they either pay full price or fall back to the previous generation.

Independent Assessment

GPT-6 Astra is a genuine technological leap. The 95% on ARC-AGI-2, the 97.6% on FrontierMath Tier 4, and a score more than double Sol's on AutomationBench—these numbers are cross-corroborated across multiple independent sources.

But this release also exposes a structural problem: as the boundaries of model capability are pushed ever further, interpretability is instead declining—OpenAI itself acknowledges this in its system card. The inverse relationship between capability and monitorability is an intrinsic contradiction of the current mainstream training paradigm, not a problem unique to Astra. But Astra is the first publicly deployed model to push this contradiction to the "Critical" level.

The 99.9% ARC-AGI-3 score will be the number remembered from this release, but it was an upper-bound result under tuned conditions. For most use cases, the 62.7% standard test score and the actual net cost difference of $40 per million tokens are the baselines that truly need to be calculated. As for Greg Brockman's declaration that the AGI era has arrived, the most honest way to read it is this: it does not mean humans can now step aside, but that AI has, for the first time, made humans optional—rather than strictly necessary—along certain task dimensions.