**GPT-6 Sol First Test: 84 in the 18-Question Targeted Evaluation**
OpenAI officially announced today that GPT-6 with Intelligent UI has been globally rolled out to Plus, Pro, Business, and Enterprise users. GPT-6 Sol has become the default model for those paid tiers, and Free and Go users will gradually receive GPT-6 Luna starting tomorrow. Within hours of the announcement, YZ Index immediately launched an 18-question targeted evaluation, with twice the number of questions of the routine Smoke test, but it is still a rapid targeted scope, not a full weekly leaderboard baseline.
The results of this 18-question targeted evaluation show that GPT-6 Sol achieved a composite score of 84. Execution was 85.7, Solidity 81.8, and the integrity rating was pass. All three core metrics were obtained within the 18-question limit, reflecting the model's immediate performance level in instruction following, content completeness, and compliance. Because the evaluation scale is only half that of the regular weekly leaderboard, these scores are for quick reference only, and the full baseline still needs to be confirmed by the next weekly Full evaluation.
When making a rough comparison with the current top ten on the main leaderboard, special attention must be paid to methodology differences. In the latest full weekly evaluation on the main leaderboard, gpt-6-sol ranks first with 82.2, claude-opus-4.7 has 81.5, gpt-o3 has 80.6, followed by gpt-5.5 (80.5), grok-4 (80.4), gpt-6-astra (80.4), claude-sonnet-4.6 (80.1), gpt-6-luna (79.4), gpt-6.1-sol (78.6), and doubao-pro (76.7). The 84 points from this 18-question targeted evaluation and the 82.2 points on the main leaderboard cannot be directly combined in the same ranking, because there are significant differences in question volume, difficulty distribution, and evaluation duration. Targeted tests focus more on immediate response and execution in specific scenarios, while the weekly Full evaluation covers broader dimensions and longer-cycle stability.
Based on known data, GPT-6 Sol's Execution of 85.7 under the 18-question scope is higher than its Solidity of 81.8, indicating that the model responds quickly in task decomposition and output structure, but there is still room for improvement in content depth and consistency. An integrity rating of pass indicates that the model showed no obvious boundary-crossing or fabricated behavior, meeting paid users' baseline safety requirements.
YZ Index will proceed with subsequent evaluations at its established pace. After the next weekly Full evaluation is completed, gpt-6-sol's full baseline score will be compared with the existing main leaderboard under the same methodology. At that time, the 18-question targeted results will serve only as preliminary reference. Users can continue to experience GPT-6 Sol's Intelligent UI features through the ChatGPT interface; actual performance remains subject to the official phased rollout.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接