GPT-6.1 Sol First Test: YZ Index 18-Question Targeted Evaluation

GPT-6.1 Sol has received its first targeted YZ Index evaluation, scoring 88 on 18 questions with an execution score of 98.7. Winzheng notes the result is n

GPT-6.1 Sol First Test: 88 Points on 18 Questions, Execution 98.7

OpenAI's official account @OpenAIDevs announced today that GPT-6.1 Sol can provide a larger workspace for coding agents, performs strongly in coding, computer use, and cross-application workflow scenarios, and costs less than GPT-6 Astra. The model has just completed an official update, and Winzheng immediately launched an 18-question targeted evaluation, with twice the number of questions of the daily Smoke test.

This evaluation uses a rapid targeted methodology, with a composite score of 88. The execution score is 98.7, demonstrating extremely high task-completion stability; Solidity is 74.8, reflecting the model's performance in detail handling and long-horizon consistency; the integrity rating passed (pass). All the above data come strictly from the 18-question targeted test, without any extrapolation or supplementation.

It should be clear that this 18-question evaluation is not on the same basis as the full weekly leaderboard, and both sample size and difficulty distribution differ from the full baseline. In the current top ten of the main leaderboard (the most recent complete weekly evaluation results), claude-opus-4.7 ranks at 83.9, grok-4 at 83.3, doubao-pro at 80.5, followed by gpt-5.5 (78.2), gpt-o3 (77.7), claude-sonnet-4.6 (76.4), gemini-3.1-pro (75.5), deepseek-v4-pro (75.2), gemini-2.5-pro (74), and qwen3-max (68.7). These scores are all based on complete weekly evaluations and cannot be directly mixed or compared with this 18-question targeted result; they are for rough reference only.

The coding and computer-use capabilities emphasized in the official announcement were somewhat corroborated in this targeted test, but the 18-question sample size is limited and cannot cover all scenarios of the weekly leaderboard. The gap between the Solidity score of 74.8 and the Execution score of 98.7 also indicates that the model still has room for optimization in different dimensions.

Winzheng will subsequently, according to the established plan, launch a complete weekly Full evaluation next week, covering a larger question set and more comprehensive dimensions, to provide a final positioning consistent with the main leaderboard methodology. Detailed sub-item breakdowns and cross-model comparisons will be released at the same time.