GPT-o3 Code Execution Plunges 24.5 Points, Main Leaderboard Falls 15.2: Random Draw or Genuine Degradation?

In today's Smoke evaluation, GPT-o3's main leaderboard score fell from 96.01 to 80.78, while its code execution dimension dropped directly from 97.00 to 72.50, a decline of 24.5 points.

Data Facts: Extreme Fluctuation in a Single Dimension

Comparing yesterday's and today's scores, code execution showed a single-day change of -24.5 points, while material constraints fell from 94.80 to 90.90, only -3.9 points. Engineering judgment remained unchanged at 100.00, and task expression rose from 75.00 to 90.00, up 15 points. The main leaderboard therefore dropped 15.2 points overall. Integrity rating remained pass.

The Smoke evaluation includes only 2 questions per dimension each day, an extremely small sample size. A 24.5-point drop in the code execution item means at least one question's score fell sharply. The 3.9-point decline in material constraints is within the normal fluctuation range.

Cause Analysis: Draw Variance Is More Likely

From the dimension composition, code execution is most sensitive to specific questions. In a 2-question test, if the model encounters questions with complex boundary conditions or requiring multi-step reasoning, a single mistake can lower the average score by 24.5 points. The smaller decline in material constraints indicates that the model's foundational ability in faithfulness has not shown systemic problems.

Engineering judgment stayed at full marks for two consecutive days, and task expression actually improved, indicating that the model's secondary leaderboard abilities did not decline in tandem. If this were genuine degradation, it would usually be accompanied by simultaneous weakening across multiple dimensions; this time only code execution stood out alone, so draw variance is a more reasonable explanation.

Implications for Users

Teams that rely heavily on code generation should note that a single-day Smoke score cannot be directly equated with everyday usage performance. A code execution score of 72.50 may only reflect the difficulty distribution of the day's 2 questions, rather than a decline in the model's overall programming ability.

Scenarios with high requirements for material faithfulness are affected to a limited extent; material constraints remain at 90.90, and the decline is not enough to change selection decisions. A full score in engineering judgment indicates that the model remains stable on judgment-type tasks.

Strategic Judgment: No Need for Excessive Concern for Now

Based on the current score comparison, GPT-o3's main leaderboard decline this time is most likely caused by question draw, rather than genuine model degradation. The stability or increase in engineering judgment and task expression further supports this judgment.

It is recommended to continue observing the code execution dimension in the next Smoke evaluation. If a similar decline of more than 24 points occurs for two consecutive days, then consider launching validation with a larger dataset. The current single-day data is insufficient to conclude that the model's ability has degraded.

Enterprises selecting models can still consider GPT-o3 as a candidate for code execution, but should pair it with multi-day average scores or larger-sample tests to avoid being misled by a single draw result.


Data source: YZ Index | Run #341 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!