A test report released by ARC Prize on September 3, 2026, shows that GPT-6 Astra scored 62.7% at a cost of $26,000 on the ARC-AGI-3 Semi-Private task set under the standard harness, and 99.9% at a cost of $19,000 under the Provider Adapter harness.
The Facts
The report draws a clear distinction between the two harnesses: the standard harness allows the model to carry notes of its own choosing throughout the environment, while the Provider Adapter harness retains opaque reasoning states and employs compression, enabling reuse of prior work across requests. Under both frameworks, the model achieved state-of-the-art results, and at the highest reasoning effort level, its action efficiency surpassed the human median, with 96% of levels completed using fewer actions than humans.
Mechanism Breakdown
The score gap stems directly from differences in state retention. The standard harness requires the model to rely solely on its own notes for each request, whereas the Provider Adapter harness allows the model's internal state to persist and compresses lengthy conversations, reducing redundant reasoning overhead. The report's table shows that as reasoning effort increases from none to max, the standard harness cost drops from $49,000 to $26,000, while the Provider Adapter harness cost drops from $23,000 to roughly $17,000.
Industry Impact
This case once again underscores how evaluation setups can determine rankings. The earlier ExploitBench perfect-score controversy had already prompted similar questions, and this episode further magnifies the systemic issue of horizontal comparability across benchmarks. Human participant testing costs approximately $12.78 per attempt, and the model's cost under efficient reasoning is now close to or below that level—yet the fairness of the two harnesses remains contested.
Strategic Assessment
[Analysis] If benchmark developers cannot unify state retention rules, vendors may increasingly choose to submit results under the harness most favorable to them, eroding the reference value of public leaderboards. Cross-vendor comparisons using the standard harness more closely reflect real-world deployment scenarios, but in the short term they may lower the ceiling for what models can demonstrate on complex tasks.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接