In today's Smoke evaluation, GPT-5.5's main leaderboard score dropped from 82.29 to 60.87, a single-day decline of 21.4 points.
Core Dimension Score Comparison
The code execution dimension fell from 97.00 yesterday to 75.00, a drop of 22 points; the material constraint dimension fell from 64.30 to 43.60, a drop of 20.7 points. Engineering judgment (side leaderboard, AI-assisted evaluation) remained unchanged at 100.00, while task expression (side leaderboard, AI-assisted evaluation) dropped from 91.70 to 81.70.
Data Facts and Cause Analysis
The Smoke evaluation has only 10 questions per day, with 2 questions per dimension, so the small sample size makes single-day fluctuations normal. This time, two main leaderboard dimensions of GPT-5.5 simultaneously showed drops of more than 20 points, far exceeding the range of typical sampling fluctuation. Both code execution and material constraint declined sharply, indicating that on the questions drawn this time, the model showed a clear drop in both the precision of instruction execution and fidelity to the given materials.
The zero fluctuation in the engineering judgment dimension shows that the model can still maintain a full-score level on side leaderboard tasks, and the task expression dimension only dropped 10 points, indicating that the problems are mainly concentrated in the auditable dimensions of the main leaderboard. The integrity rating remains pass, and no threshold warning was triggered.
Implications for Users
Teams that rely heavily on code execution should immediately reduce their dependence on GPT-5.5, especially in scenarios requiring high-precision script generation or complex logic checking. Risks rise in scenarios sensitive to material constraints (such as contract review, code review, and knowledge base question answering); a score of 43.60 means the model is more likely to deviate from the given materials.
Engineering judgment maintaining a full score still has reference value for developers who need decision support, but the main leaderboard collapse has limited overall usability.
Strategic Assessment
This 21.4-point drop has exceeded the normal fluctuation range of the Smoke evaluation and warrants close attention in the next period. If the score remains around 60 the next day, GPT-5.5's main leaderboard capability may have genuinely degraded; if it rebounds quickly, this was a typical small-sample draw event. Current data does not support continuing to list GPT-5.5 as a primary candidate model.
The main leaderboard consists only of code execution and material constraint; simultaneous failure in both dimensions directly caused the overall ranking decline. The relative stability of engineering judgment and task expression cannot make up for the main leaderboard gap.
The 60.87-point main leaderboard score, formed by 75.00 in code execution and 43.60 in material constraint, is already below the daily level of most comparable models.
Enterprises selecting models should prioritize observing the results of the next three Smoke runs before deciding whether to adjust deployment strategy. Developers relying on GPT-5.5 need to prepare alternative models to cope with a possible continued low main leaderboard position.
Data source: YZ Index | Run #320 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接