GPT-o3 Code Execution Plunges 32.5 Points as Main Leaderboard Falls from 96.01 to 80.48

In today's Smoke evaluation, GPT-o3's main leaderboard score fell from 96.01 yesterday to 80.48, a drop of 15.5 points, with the code execution dimension plunging from 97.00 to 64.50, a decline of 32.5 points.

Breaking Down the Data: Changes in Core Dimensions

The main leaderboard consists of only two auditable dimensions: code execution and material constraints. Today, code execution scored 64.50 and material constraints scored 100.00, yielding a simple average of 80.48. Yesterday's corresponding figures were 97.00 for code execution and 94.80 for material constraints, averaging 96.01. Material constraints actually improved by 5.2 points today, indicating no systemic problem with the model's fidelity to source material.

Data from the side leaderboard (AI-assisted evaluation) shows engineering judgment falling from 88.90 to 77.30, while task expression rose from 65.00 to 90.00. The integrity rating remained at pass, with no threshold warnings triggered.

Cause Analysis: Sampling Noise or Real Degradation

The Smoke evaluation uses only 2 questions per dimension per day, 10 questions in total, an extremely small sample. A single-day 32.5-point drop in code execution is most likely due to differences in difficulty or question type introduced by random sampling, rather than parameter-level degradation of the model. The simultaneous 5.2-point rise in material constraints further supports the judgment that this is fluctuation rather than an overall capability decline. The 11.6-point drop in engineering judgment may be linked to the code execution questions, but the 25-point rise in task expression shows the model maintains or improves response quality on expression-type tasks.

If this were genuine degradation, material constraints should not have improved at the same time; the current opposing movement across the two dimensions is more consistent with random fluctuation in a small sample.

Implications for Users

Teams that rely heavily on code execution should add redundant verification steps when deploying GPT-o3, especially in mathematical computation and algorithm implementation scenarios. A material constraints score of 100.00 indicates the model remains suitable for document processing tasks that require strict adherence to source text. An engineering judgment score of 77.30 (side leaderboard, AI-assisted evaluation) suggests that complex system design work requires human review.

Developers who base model selection decisions on a single day's Smoke results may overestimate or underestimate true stability.

Strategic Assessment

Based on today's data, the consistency of GPT-o3's code execution warrants continued tracking in the next period. The main leaderboard score of 80.48 remains above most competitors' baselines, but a 32.5-point swing in a single dimension exceeds the normal range. The coexistence of a perfect material constraints score and a low code execution score shows the model performs unevenly across different capability axes. We recommend that the next Smoke round focus on whether code execution rebounds above 90 points, in order to distinguish incidental fluctuation from systematic change.

There is currently no evidence supporting overall model degradation; users can continue to observe rather than switch immediately.


Data source: YZ Index | Run #335 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!