GPT-5.5 Code Execution Falls from 100 to 75; Main Leaderboard Drops 10.1 Points in One Day

In today's Smoke evaluation, GPT-5.5's main leaderboard score fell from 86.50 to 76.44, a decline of 10.1 points.

Score Comparison Data

The Code Execution dimension scored 100.00 yesterday and 75.00 today, a decline of 25 points. The Material Constraints dimension scored 70.00 yesterday and 78.20 today, an increase of 8.2 points. Engineering Judgment (side leaderboard, AI-assisted evaluation) rose from 50.00 to 100.00. Task Expression (side leaderboard, AI-assisted evaluation) edged down from 91.70 to 90.00. The Integrity rating remained pass.

Data Fact Breakdown

The main leaderboard consists of two auditable dimensions: Code Execution and Material Constraints. Today's main leaderboard score of 76.44 comes directly from the weighted average of Code Execution at 75.00 and Material Constraints at 78.20. The 25-point drop in the single Code Execution dimension is the sole source of the main leaderboard decline; the 8.2-point rebound in Material Constraints only partially offset that loss.

Cause Analysis

The Smoke evaluation uses only 2 questions per dimension each day, so the sample size is extremely small. The most likely reason Code Execution fell from 100.00 to 75.00 is that the two questions drawn today were more difficult than yesterday's, causing a large fluctuation in the model's score on that dimension. The 8.2-point increase in Material Constraints likewise points to differences in question sampling rather than a systemic change in model capability. Engineering Judgment rising from 50.00 to 100.00 further confirms the impact of single-day question randomness on the side leaderboard.

Real degradation versus question fluctuation must be distinguished using multiple consecutive days of data. With only a single-day comparison, it is impossible to confirm whether the model's Code Execution capability has undergone a structural decline. The Integrity rating remained pass, indicating no integrity issues in the model's responses.

Implications for Users

Teams that rely heavily on Code Execution should note that GPT-5.5's single-day score on this dimension can fluctuate by as much as 25 points. Developers who depend on the model to generate reliable code should add a manual review step and avoid treating today's 75.00 score as the norm. The rise in Material Constraints to 78.20 provides some buffer for scenarios requiring the model to strictly follow input constraints.

Engineering Judgment rising to 100.00 is a positive signal for scenarios requiring the model to assist with engineering decisions, but this dimension is on the side leaderboard, and AI-assisted evaluation results are for reference only.

Strategic Assessment

This 10.1-point decline in the main leaderboard was mainly driven by the 25-point drop in the Code Execution dimension. Combined with the fact that the Smoke evaluation has only 10 questions per day, the change is more likely caused by question sampling fluctuation rather than real model degradation. The next evaluation should focus on whether the Code Execution dimension rebounds above 90. If Code Execution scores below 80 on two consecutive days, that signal should be flagged as a potential capability change.

The combination of Material Constraints at 78.20 and Engineering Judgment at 100.00 shows that GPT-5.5 remains usable for constraint adherence and engineering judgment. Enterprises selecting models can continue to keep GPT-5.5 as one part of a multi-model portfolio, but should reduce its weight in pure code generation tasks.


Data source: YZ Index | Run #332 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!