GPT-o3 Code Execution Plummets 21.8 Points; Material Constraints Rise 33.5 Points; Main Leaderboard Edges Up

In today's Smoke evaluation, GPT-o3's code execution score fell from 93.80 yesterday to 72.00, a drop of 21.8 points, while its material constraints score rose from 49.30 to 82.80, an increase of 33.5 points, and its main leaderboard score rose from 73.78 to 76.86.

Breaking Down the Data Facts

This evaluation covers only 10 questions per day (2 per dimension). The code execution dimension's 93.80 yesterday corresponded to a relatively high accuracy rate; today's 72.00 means at least 1 of the 2 questions had a clear error. The material constraints dimension's 49.30 yesterday corresponded to a relatively low constraint-following rate; today's 82.80 indicates that both questions reasonably satisfied the material limits. Engineering judgment rose from 88.90 to 100.00 (side leaderboard, AI-assisted assessment), task expression rose from 66.70 to 90.00 (side leaderboard, AI-assisted assessment), and the integrity rating changed from warn to pass.

Cause Analysis: Random Draw Volatility or Real Degradation

The Smoke evaluation's single-day sample size is extremely small; just 2 questions can cause swings of more than 20 points. The simultaneous sharp, opposite changes in code execution and material constraints are most likely explained by the difficulty or trap distribution of today's drawn code execution questions differing from yesterday's, while the material constraints questions happened to match the model's current strengths. The available data does not show same-dimension declines across multiple consecutive days, so it cannot support a conclusion of real model degradation.

Specific Implications for Users

Teams that rely heavily on code execution should add local verification steps when calling GPT-o3, because a single-day score of 72.00 means the model's error rate has risen significantly in some code scenarios. Material-constraint-sensitive scenarios (such as long-document rewriting and strictly formatted output) improved today, and 82.80 can serve as a short-term reference value. The main leaderboard score of 76.86 remains in the midrange; for mixed tasks that require both code and material constraints, the model's stability has not yet reached a level that can go directly into production.

Strategic Judgment

The combination of this 21.8-point drop and 33.5-point rise is more likely random volatility caused by question drawing rather than a structural change in model capability. It is recommended that the next period continue tracking the code execution dimension, and if it stays below 80 for two consecutive days, launch deep retesting; the current data does not support downgrading the long-term assessment of GPT-o3's code execution ability.


Data source: YZ Index | Run #345 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!