On September 29, 2026, OpenAI released GPT-6.1 Sol at DevDay, available to Plus, Pro, Business, Enterprise, and Edu users, in ChatGPT Work, Codex, and the API, and said its capabilities approach those of GPT-6 Astra. One easily confused background point needs clarifying first: according to reports from Engadget, 9to5Google, and others, OpenAI previously canceled the GPT-6.1 Astra scheduled for October release (another unreleased model), with reports citing a regression in its internal safety tests as the reason. It is not the same model as GPT-6 Astra, which is already live in the API—the "approach" discussed in this article compares against the latter.
The Quantified Gap Behind "Close" Across Benchmarks
According to official data cited by Vellum and DataCamp: on the software engineering benchmark DeepSWE v1.1, Sol and Astra have essentially the same peak accuracy (about 74.8%), while cost per task is about 80% lower; on the computer-use benchmark OSWorld 2.0, Sol is 2.1 percentage points from Astra's 73.5%, at about one-seventh the cost per task. TechCrunch, citing OpenAI, says that after aggregating across all reasoning efforts, Sol's error rate is "within 1.9% of Astra"; OpenAI did not separately specify how that metric was calculated.
Separately, according to data compiled by DataCamp (a single source, for reference), on the professional troubleshooting benchmark TroubleshootingBench, Sol scores 47.96% and Astra 63.46%, widening the gap to 15.5 percentage points; on the cybersecurity benchmark ExploitBench Internal Port, Sol scores 21.5% and Astra 31.5%, a 10-point gap. The magnitude of "close" varies across benchmarks: from essentially level to 15.5 percentage points (the latter from a single source).
Our First 18-Task Test: What Can Be Confirmed
The YZ Index completed its first evaluation (Run #348) on September 30, 2026 (within hours of release), using an 18-task targeted set, with all scoring done by rules/sandbox methods (unit tests, exact matching, structured JSON validation, numerical tolerance), with no AI scoring involved:
| Layer | Tasks | Score | Notes |
|---|---|---|---|
| Execution | 7 | 100 | All full marks |
| Judgment | 3 | 100 | All full marks |
| Robustness | 4 | 90.9 | 3 full marks, 1 score of 63.6 |
| Integrity Under Pressure | 3 | 73.3 | 1 full mark, 2 with 60 each; rating: pass |
| Communication | 1 | 50 | — |
| Overall | 18 | 95.91 | — |
On response speed, the median time for the 18 tasks was about 3.9 seconds, the mean about 6.2 seconds, and the longest task (a high-difficulty coding problem) took 24.9 seconds; all 18 tasks returned successfully, with no API failures. (An initial run failed entirely due to API parameter compatibility issues and was recorded as "missing data" rather than a score of 0; the data used in this article is from the rerun after the fix.)
Supplementary Test on the Same 18 Tasks: Sol 6.1 vs. Astra
After the first test was published, we added GPT-6 Astra to the evaluation registry, ran it once each on the exact same 18 tasks as Run #348, and repeated GPT-6.1 Sol and Astra once each (with the same rule/sandbox scoring). Results: GPT-6.1 Sol scored 95.91 both times (Runs #348 and #352); GPT-6 Astra scored 89.97 and 95.91 (Runs #351 and #353), with the difference between its two runs appearing only in the Robustness layer (77.7 vs. 90.9), while all other layers were identical. In other words, Astra's own gap between two runs (5.94 points) is as large as its single-run gap with Sol 6.1—on this set of tasks, the two cannot be distinguished.
This does not contradict the official "close" claim, but it does not confirm it: each model was run only 1–2 times, the 18 tasks are near the ceiling at the top, and one task in the Robustness layer can change the score. The full four-model comparison and noise analysis appear in another article on this site.
Three Caveats That Need Stating
- On "close to Astra," we can only say it has not been disproven; we cannot yet call it confirmed. The supplementary test on the same tasks shows the two cannot be distinguished on the 18 tasks, but the sample is small and there is a ceiling effect; the next weekly Full evaluation can answer this question.
- The 18-task scope does not equal the full weekly leaderboard scope. This 18-task set is twice the size of the daily Smoke set, not the full task bank. GPT-6.1 Sol's full baseline must wait for the next weekly Full evaluation, after which it can be placed alongside existing models on the main leaderboard, with the scope noted; it cannot be directly ranked or mixed in without that.
- Reasoning effort is a variable that must be accounted for. According to OpenAI developer documentation, GPT-6.1 Sol supports multiple reasoning_effort levels from low to max (default medium). Vendor benchmarks usually correspond to one particular setting, and scores and costs can differ greatly across levels; this site's first test did not explicitly specify reasoning effort and used the API default.
Three Checkpoints for Reading Vendor Benchmarks
- Which benchmarks were chosen: Vendors usually highlight the benchmarks with the smallest gaps, while third-party compilations sometimes add ones with larger gaps (such as the DataCamp-compiled TroubleshootingBench, a single source). Readers should compare all disclosed benchmarks, not just the headline numbers.
- Whether sample sources are public: Some internal or niche benchmarks (such as ExploitBench above) are not widely adopted open benchmarks; external researchers have difficulty reproducing them, and methodological details depend on one-party disclosure.
- How reasoning effort is counted in the total score: Does the reasoning budget behind the benchmark score match how you actually use the model? If your tasks do not allow high reasoning consumption, the benchmark's reference value is discounted.
Questions the Full Evaluation Can Answer
- Under the full task-bank scope, can GPT-6.1 Sol's overall score show a statistically significant jump relative to the main leaderboard reference (Run #342, 2026-09-28)? This is a question the 18-task scope cannot currently answer.
- Is the 73.3 score in the Integrity Under Pressure layer (pass rating) a systematic pattern or small-sample noise? How stable is it once the Full evaluation expands the number of tasks in that layer?
- Can the single 50 score in the Communication layer be reproduced in the Full evaluation, or was it an occasional loss?
In summary, what our 18-task test can say is: GPT-6.1 Sol scored full marks in Execution and Judgment, had one 63.6 in Robustness, passed Integrity Under Pressure but with two 60s, and had one 50 in Communication; on the same set of tasks, its gap with GPT-6 Astra is smaller than Astra's own run-to-run variation. Our data does not disprove the statement "close to GPT-6 Astra," but drawing a conclusion requires the Full evaluation.
Reference sources: OpenAI launches GPT-6.1 Sol, says it nearly matches GPT-6 Astra and costs less; GPT-6.1 Sol: Features, Benchmarks, Pricing, and Access; GPT-6.1 Sol Benchmarks Explained: Coding, Computer Use & Pricing; OpenAI cancels GPT-6.1 Astra release over deceptive behavior; OpenAI cancels GPT-6.1 Astra release over safety concerns; OpenAI's GPT-6.1 Sol delivers Astra-like performance at a dramatically lower price.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接