GPT-5.5's main leaderboard score in today's Smoke evaluation dropped from 95.01 to 85.93, a decline of 9.1 points.
Score Breakdown and Key Drivers
Code execution fell from 100.00 to 97.00, down 3 points; material constraints fell from 88.90 to 72.40, down 16.5 points. Engineering judgment rose from 75.00 to 100.00, up 25 points; task expression remained unchanged at 90.00. The main leaderboard includes only the two dimensions of code execution and material constraints, and the 16.5-point drop in material constraints directly caused a net loss of 9.1 points on the main leaderboard.
Question Draw Fluctuation or Genuine Model Degradation?
The Smoke evaluation uses only 10 questions per day, 2 per dimension, so single-day fluctuation is normal. The concentrated losses in material constraints, alongside a sharp rise in engineering judgment, indicate that the material constraint questions drawn that day may have been harder with stricter constraints, while the engineering judgment questions were relatively looser. This kind of reverse fluctuation across dimensions is more consistent with the randomness of question draw than with overall model capability degradation. Code execution fell only 3 points, a limited magnitude, further supporting the fluctuation explanation.
Practical Implications for Users
Scenarios that heavily depend on material constraints (such as long-document fact checking, citation generation, and enterprise knowledge base Q&A) require heightened vigilance — today's 72.40 is already below yesterday's 88.90, showing a clear decline in single-day usability. Teams that rely on code execution are less affected, as 97.00 remains at a high level. Although the engineering judgment score of 100.00 is high, it belongs to the side leaderboard (AI-assisted evaluation) and does not directly affect the main leaderboard ranking.
Strategic Assessment
The 16.5-point single-day drop in material constraints has exceeded the normal fluctuation threshold and warrants continued tracking in the next round. If material constraints recover to above 85 points in the next evaluation, it can be attributed to draw fluctuation; if it remains below 80 points, then a genuine decline in the model's stability on this dimension must be considered. Current data does not support the conclusion of "overall model degradation," but rather "abnormal fluctuation in the material constraints dimension."
The integrity rating remains "pass," with no threshold entry issues triggered. The stability dimension was not provided in this round's data, so consistency cannot be directly assessed.
Data source: YZ Index | Run #298 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接