GPT-5.5 Code Execution Drops Sharply from 100 to 75; Smoke Evaluation Main Leaderboard Falls 7.5 Points

In today's Smoke evaluation, GPT-5.5's code execution score fell from 100.00 to 75.00, and its overall main leaderboard score slipped from 79.12 to 71.67.

Score Change Facts

The code execution dimension dropped 25 points; material constraints rose from 53.60 to 67.60; engineering judgment fell from 100.00 to 75.00; task expression climbed from 81.70 to 90.00; and the integrity rating remained pass. The main leaderboard is derived from a weighted combination of code execution and material constraints, resulting in a net decline of 7.5 points today.

Analysis of Possible Causes

The Smoke evaluation uses only 10 questions per day, with 2 questions per dimension, so each question carries high scoring weight. The simultaneous 25-point drops in code execution and engineering judgment suggest that the day's sampled questions placed greater demands on code correctness or logical consistency. The 14-point rise in material constraints, meanwhile, indicates an improvement in the model's performance under strict citation restrictions—pointing not to an overall capability regression but to the impact of question-bank sampling on specific dimensions.

If this were a genuine model regression, it would typically be accompanied by simultaneous declines across multiple dimensions and appear over consecutive days. With only a single day of data, no regression trend can yet be confirmed. Random sampling fluctuation is more likely, but a 25-point drop does exceed the range of normal day-to-day variation.

Implications for Users

Development teams that rely heavily on code execution should be alert to volatile single-day performance. A code execution score of 75 means GPT-5.5 has a higher probability of errors in scenarios involving complex logic or multi-step computation. A material constraints score of 67.60 indicates that the model still has clear limitations in strict citation tasks, making it better suited to content generation scenarios where fidelity requirements are lower.

The simultaneous 25-point drop in engineering judgment poses a direct risk to projects that require the model to assist with architecture decisions or code review. The 8.3-point rise in task expression has limited impact on pure text-output tasks.

Strategic Assessment

This change is most likely caused by question-selection fluctuation, but the magnitude of the drop has reached a threshold that warrants continued tracking. If code execution still scores below 85 in the next Smoke evaluation round, it would be reasonable to consider whether the model has a systemic problem in this dimension. Current data does not support the conclusion that the model has degraded; it only supports a finding of significant single-day fluctuation.

Enterprises in the model selection process should add GPT-5.5's code execution performance to their daily monitoring list rather than immediately replacing the model. The improvement in material constraints warrants further validation to determine whether it is a stable gain.


Data source: YZ Index | Run #308 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!