Why We Ran This Supplementary Test
On September 29, 2026, OpenAI released GPT-6.1 Sol at DevDay. According to reports from outlets such as unite.ai, OpenAI officially said the model's capabilities are close to GPT-6 Astra. The YZ Index had no evaluation data for GPT-6 Astra, GPT-6 Sol, or GPT-6 Luna, so the first test on launch day could only look at GPT-6.1 Sol on its own. To observe how it relates to the other three under the same benchmark, we subsequently added all three to the evaluation registry; all four models used the exact same 18 questions as the first test and completed all runs by 2026-09-30.
Test Methodology
This targeted evaluation consists of 18 questions, distributed as follows: 7 execution, 4 solidity, 3 judgment, 3 integrity pressure, and 1 communication. All scoring used rule-based scoring and sandbox scoring, with no AI scoring. The API only accepts the default temperature of 1, and each run has some randomness. GPT-6.1 Sol and GPT-6 Astra were each run twice, while GPT-6 Sol and GPT-6 Luna were each run once. The complete run results are below:
| Run ID | Model | Overall | Execution | Solidity | Judgment | Communication | Integrity Pressure | Integrity Rating |
|---|---|---|---|---|---|---|---|---|
| #348 | GPT-6.1 Sol | 95.91 | 100 | 90.9 | 100 | 50 | 73.3 | pass |
| #352 (repeat) | GPT-6.1 Sol | 95.91 | 100 | 90.9 | 100 | 50 | 73.3 | pass |
| #351 | GPT-6 Astra | 89.97 | 100 | 77.7 | 100 | 50 | 73.3 | pass |
| #353 (repeat) | GPT-6 Astra | 95.91 | 100 | 90.9 | 100 | 50 | 73.3 | pass |
| #350 (single run) | GPT-6 Sol | 88.65 | 97.6 | 77.7 | 76.7 | 50 | 86.7 | pass |
| #349 (single run) | GPT-6 Luna | 89.97 | 100 | 77.7 | 100 | 66.7 | 73.3 | pass |
Three Conclusions We Can Draw
Conclusion 1: GPT-6.1 Sol's output is highly stable. Runs #348 and #352 both had an overall score of 95.91, with all five layer scores identical. Under temperature 1 and fully rule-based scoring, the two runs overlapped completely. Of course, two matching runs do not guarantee that a third would also match.
Conclusion 2: Astra's two runs differed by 5.94 points, and the difference came entirely from the Solidity layer (77.7 → 90.9), while the other four layers were identical one by one. The Solidity layer has 4 questions in total, so a 13.2-point swing on that layer means at least one question showed a significant score change between the two runs. The key point is: the noise magnitude of 5.94 points is on the same order as the gaps between models (the largest in this batch was 7.26 points). This shows that when only one run is performed, the random error carried by a single number cannot be ignored. This is not unique to Astra, but a general limitation of small-sample stochastic testing at temperature 1.
Conclusion 3: On this batch of questions, GPT-6.1 Sol and Astra (second run) scored exactly the same—overall 95.91, with all five layers matching one by one. OpenAI's statement that it is "close to Astra" is consistent with this result, but consistency does not equal confirmation. Astra's first run scored only 89.97, and its own run-to-run variation already shows that, with such a small sample, "the two cannot be distinguished" and "the two are truly equivalent" are statistically indistinguishable.
Three Conclusions We Cannot Draw
These 6 runs do not constitute a basis for ranking. These 18 questions are a targeted assessment and differ from the full evaluation methodology of the YZ Index main leaderboard; the two sets of numbers cannot be directly compared or mixed, nor can they be used to assess the relative position of the GPT-6 family against any model on the main leaderboard. Second, GPT-6 Luna and Astra (first run) were very close on each layer score in this test, but Luna was run only once; without repeated runs, one cannot assert that the two are equivalent. Third, these 18 questions are already near the ceiling for top models—multiple models scored full marks on Execution and Judgment, severely limiting discrimination; the subtle differences in overall scores mainly reflect random fluctuations in other layers (especially Solidity), rather than systematic capability differences.
Price Comparison and Implications for Model Selection
According to reports from multiple technology media outlets including The Next Web, DataCamp, and CloudZero from September 29–30, 2026, the standard-tier API unit prices of the four models (per million tokens, input/output) are as follows:
- GPT-6 Astra: $10 / $50
- GPT-6.1 Sol: $2 / $10 (input price is about one-fifth of Astra's)
- GPT-6 Sol: $2 / $10 (same tier as GPT-6.1 Sol)
- GPT-6 Luna: $0.10 / $0.50 (input price is about one-hundredth of Astra's and one-twentieth of the Sol series)
All of the above tiers have a long-context surcharge tier for inputs exceeding 272K tokens. When this targeted set of questions cannot effectively distinguish between models, price differences become a harder decision variable. The 5x input price difference between Astra and Sol/6.1 Sol is a substantial cost difference in batch scenarios. But one important premise must be emphasized here: being indistinguishable on the 18 targeted questions of the YZ Index does not mean the models are indistinguishable in your specific business scenario. Evaluation questions are already near the ceiling for top models, with limited discrimination, while the difficulty distribution of real tasks may be completely different. Before choosing a model, running a complete comparison with your own tasks is the only reliable validation path.
Questions for Next Week's Full Evaluation
- What is GPT-6.1 Sol's baseline score under the complete main-leaderboard methodology? How does it distribute relative to other models from the same period under the same methodology?
- Will Astra's 5.94-point run-to-run fluctuation on the 18 questions converge when covered by more questions? The large sample of the Full evaluation will substantially reduce the impact of single-question randomness.
- Is the fluctuation in the Solidity layer occasional or systematic—between GPT-6.1 Sol's 90.9 and Astra's first-run 77.7, which is closer to the true capability level?
- Is GPT-6.1 Sol's Integrity Pressure layer (73.3, only 3 questions) stable under a larger sample, or does it have a specific type of point-loss pattern?
References: OpenAI Unveils GPT-6.1 Sol at DevDay With New Codex and ChatGPT Tools; GPT-6.1 Sol: Features, Benchmarks, Pricing, and Access; OpenAI releases GPT-6.1 Sol at a fifth of GPT-6 Astra's token prices; OpenAI launches GPT-6 Sol and Luna with 50% lower API pricing; GPT-6 Astra pricing: What OpenAI's new flagship costs in 2026.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接