GPT-6.1 Sol First Test: 73 Overall on the 18-Question Protocol
OpenAI's official @OpenAIDevs account announced today that Ultrafast is rolling out GPT-6.1 Sol across the API, Codex, and ChatGPT Work. Within hours of the new model's release, Winzheng completed a targeted evaluation.
This first test used a rapid 18-question targeted protocol, twice the volume of a routine Smoke test and not the full weekly leaderboard protocol. Results show GPT-6.1 Sol with an overall score of 73, Execution of 73.2, Robustness of 72.7, and an Integrity rating of pass. All four figures come from this 18-question targeted evaluation, with no extrapolation or adjustment.
Looking at the score distribution, Execution at 73.2 and Robustness at 72.7 differ by only 0.5 points, indicating the model is fairly balanced between task completion and content reliability. The Integrity rating of pass confirms that the model showed no violations or runaway hallucination in this evaluation. The overall score of 73 is the direct score under the 18-question protocol; the full baseline will be determined by the next weekly Full evaluation.
The current main leaderboard shows the top 10 from the most recent full weekly evaluation. Its protocol differs markedly from this 18-question targeted test, and the two cannot be directly mixed. On the main leaderboard, gpt-6-sol ranks first with 82.2 overall; claude-opus-4.7 has 81.5; gpt-o3 has 80.6; gpt-5.5 has 80.5; grok-4 and gpt-6-astra are tied at 80.4; claude-sonnet-4.6 has 80.1; gpt-6-luna has 79.4; gpt-6.1-sol has 78.6; and doubao-pro has 76.7. The main leaderboard figures are for reference only: the 18-question protocol used for GPT-6.1 Sol's first test differs from the main leaderboard's full weekly evaluation in question count, difficulty distribution, and evaluation duration.
Winzheng emphasizes that this first test was a targeted evaluation made in rapid response to the official update, aimed at capturing the model's basic performance at the earliest opportunity. A complete baseline test of GPT-6.1 Sol will follow under the weekly Full evaluation standard, covering more questions and a stricter protocol to provide a full score directly comparable with the main leaderboard. Users can follow Winzheng's subsequent weekly leaderboard updates for the final, complete data.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接