In today's Smoke evaluation, GPT-5.5's code execution score dropped from 72.00 to 50.00, its material constraints score rose from 69.10 to 95.00, and its overall leaderboard score changed from 70.70 to 70.25.
Specific Changes in Dimension Scores
The code execution dimension fell by 22 points in a single day, the material constraints dimension rose by 25.9 points in a single day, the task expression dimension dropped by 8.3 points, and the engineering judgment dimension remained unchanged at 100.00. The overall leaderboard is a weighted combination of code execution and material constraints, and ultimately fell by only 0.5 points.
Analysis of the Causes of Fluctuation
The Smoke evaluation has only 10 questions per day, with 2 questions corresponding to one dimension. The simultaneous opposite movements of more than 20 points in the two core dimensions, code execution and material constraints, most likely stem from the randomness of question sampling. The code execution questions may have sampled cases requiring multi-step debugging or complex API calls, while the material constraints questions sampled cases requiring strict adherence to prompt boundaries. The smaller score changes in engineering judgment and task expression also support the conclusion that this fluctuation is mainly concentrated in the two sampling-sensitive dimensions.
If the model had experienced genuine degradation, it would usually affect multiple dimensions at the same time or cause a larger drop in the overall leaderboard. This time, the overall leaderboard fell by only 0.5 points, indicating that the score changes in the two dimensions largely offset each other after weighting, and there is not yet evidence of a systemic decline in capability.
Practical Implications for Users
Teams that rely heavily on code execution should, today and in the short term, additionally verify the correctness and executability of code generated by GPT-5.5 when using it. The positive impact of the sharp rise in the material constraints score is mainly reflected in scenarios that require strict adherence to format, length, or citation rules, such as contract clause generation or formatted report output.
For mixed tasks requiring both code and material constraints, GPT-5.5's performance today is close to yesterday's level, and its overall leaderboard score of 70.25 remains within the usable range. Developers can continue to use it, but for code execution results, it is recommended to add manual or automated testing steps.
Strategic Judgment
The most direct signal from this data is that single-day fluctuations in the Smoke evaluation have limited impact on the overall leaderboard. The combination of a code execution score of 50.00 and a material constraints score of 95.00 still maintains an overall leaderboard score of 70.25 after weighting, indicating that the two dimensions are highly complementary. In the next evaluation, it will be necessary to observe whether code execution rebounds above 65 points, or whether material constraints falls below 80 points, in order to determine whether this was an isolated sampling event.
Engineering judgment remains at full score and the integrity rating remains pass, indicating that there are no problems with the model's foundational capabilities or output standards. The current data does not support a downward revision of judgment on GPT-5.5's overall capability; only short-term observation of the code execution dimension is needed.
Data source: YZ Index | Run #354 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接