In today's Smoke evaluation, GPT-5.5 saw its code execution score drop from 94.50 to 72.00, a decrease of 22.5 points, while its material constraint score rose from 71.40 to 100.00, an increase of 28.6 points. The main leaderboard score only moved from 84.11 to 84.60.
Data Facts: Two Dimensions Move Sharply in Opposite Directions
The Smoke evaluation uses only 2 questions per dimension per day. The code execution dimension dropped from 94.50 yesterday to 72.00 today, while the material constraint dimension rose from 71.40 yesterday to 100.00 today. Engineering judgment remained unchanged at 100.00, and task expression rose from 81.70 to 91.70. The integrity rating remained pass. The weighted main leaderboard result rose by only 0.5 points.
Cause Analysis: Random Question Sampling Is More Likely
The single-day 22.5-point drop in the code execution dimension is most likely attributable to at least one of the two questions involving a sudden change in difficulty or type. The material constraint dimension simultaneously rose by 28.6 points to a perfect score, indicating that the model drew easier scoring questions in this round of constraint-following items. Two dimensions moving to extremes at the same time and in opposite directions points to randomness in question sampling rather than model-level parameter degradation.
If this were genuine degradation, it would typically be accompanied by multiple dimensions declining together or a notable drop in the main leaderboard. This time the main leaderboard rose by only +0.5, and engineering judgment showed zero change, further supporting the fluctuation explanation.
Implications for Users
Teams that rely heavily on code execution need to add manual verification steps to their production workflows, especially when Smoke evaluation shows a single-dimension swing of more than 20 points. A perfect material constraint score is favorable for scenarios that require strict adherence to instructions, but it cannot offset the risk posed by unstable code execution.
When selecting models, developers should factor in the historical standard deviation of the code execution dimension rather than looking only at a single day's score.
Strategic Assessment
This change is most likely caused by random question sampling, and there is currently no basis for downgrading the overall capability assessment of GPT-5.5. If code execution rebounds above 90 in the next Smoke evaluation, this can be confirmed as an isolated fluctuation; if it remains below 80, a longer-cycle stability tracking exercise should be initiated.
The current data does not support a "model degradation" judgment. It is recommended to treat GPT-5.5's code execution dimension as a high-volatility item and keep human fallback in place for critical tasks.
Data source: YZ Index | Run #317 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接