In today's Smoke evaluation, GPT-5.5's main leaderboard score fell from 91.32 to 78.80, a decline of 12.5 points, driven primarily by the code execution dimension dropping from 100.00 to 72.00.
Breakdown of Score Changes
The code execution dimension fell 28 points in a single day, the material constraints dimension rose from 80.70 to 87.10, the engineering judgment dimension held steady at 100.00, and the task expression dimension climbed from 87.50 to 100.00. The integrity rating remained "pass." The Smoke evaluation includes only 2 questions per dimension each day, 10 in total, so single-day fluctuations fall within the normal range.
Analysis of Possible Causes
The scores on the two code execution questions dragged down the overall result, while the material constraints and task expression dimensions improved, indicating that the model's performance varies markedly across question types. Randomness in question selection is the primary explanation: when the two code execution questions involve complex multi-step reasoning or edge cases, the model's output is prone to execution errors, causing a sharp drop in that dimension's score. The rise in the material constraints dimension shows that the model's performance on citation constraints did not degrade in tandem.
A genuine capability regression is unlikely. The engineering judgment dimension held at a perfect score and the task expression dimension gained 12.5 points, indicating no systemic decline in the model's secondary-leaderboard capabilities. A genuine regression would typically involve simultaneous declines across multiple dimensions, rather than an extreme swing in a single dimension.
Implications for Users
Development teams that rely heavily on code execution should be on heightened alert. Losing points on the two code execution questions in the Smoke evaluation means that the model's output stability may decline intermittently when generating runnable scripts or handling multi-file interaction scenarios. The rebound in the material constraints dimension has relatively little impact on document generation scenarios that require strict source citation.
For enterprises evaluating models, GPT-5.5's sustained perfect score in the engineering judgment dimension makes it well suited to tasks requiring highly consistent judgment. With the task expression dimension rising to 100.00, it is also suited to scenarios that generate structured output or multi-turn dialogue.
Strategic Assessment
This main-leaderboard decline mainly reflects single-day question-draw fluctuation rather than a systemic degradation of model capability. The 28-point drop in the code execution dimension exceeds the normal fluctuation range, so it is advisable to continue tracking that dimension's score in the next Smoke evaluation. If the code execution score falls below 85 for two consecutive days, a longer-cycle evaluation of more than 10 questions should be launched to verify true stability.
The current data does not support downgrading the assessment of GPT-5.5's overall capability. The simultaneous rise in the material constraints and task expression dimensions shows the model still has advantages in constraint adherence and clarity of expression. Developers can continue using it, but should add a human review step for code-execution-intensive scenarios.
Data source: YZ Index | Run #340 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接