In today's Smoke evaluation, GPT-o3's code execution score fell from 93.80 yesterday to 72.00, a drop of 21.8 points, while its material constraints score rose from 49.30 to 82.80, an increase of 33.5 points, and its main leaderboard score rose from 73.78 to 76.86.
Breaking Down the Data Facts
This evaluation covers only 10 questions per day (2 per dimension). The code execution dimension's 93.80 yesterday corresponded to a relatively high accuracy rate; today's 72.00 means at least 1 of the 2 questions had a clear error. The material constraints dimension's 49.30 yesterday corresponded to a relatively low constraint-following rate; today's 82.80 indicates that both questions reasonably satisfied the material limits. Engineering judgment rose from 88.90 to 100.00 (side leaderboard, AI-assisted assessment), task expression rose from 66.70 to 90.00 (side leaderboard, AI-assisted assessment), and the integrity rating changed from warn to pass.
Cause Analysis: Random Draw Volatility or Real Degradation
The Smoke evaluation's single-day sample size is extremely small; just 2 questions can cause swings of more than 20 points. The simultaneous sharp, opposite changes in code execution and material constraints are most likely explained by the difficulty or trap distribution of today's drawn code execution questions differing from yesterday's, while the material constraints questions happened to match the model's current strengths. The available data does not show same-dimension declines across multiple consecutive days, so it cannot support a conclusion of real model degradation.
Specific Implications for Users
Teams that rely heavily on code execution should add local verification steps when calling GPT-o3, because a single-day score of 72.00 means the model's error rate has risen significantly in some code scenarios. Material-constraint-sensitive scenarios (such as long-document rewriting and strictly formatted output) improved today, and 82.80 can serve as a short-term reference value. The main leaderboard score of 76.86 remains in the midrange; for mixed tasks that require both code and material constraints, the model's stability has not yet reached a level that can go directly into production.
Strategic Judgment
The combination of this 21.8-point drop and 33.5-point rise is more likely random volatility caused by question drawing rather than a structural change in model capability. It is recommended that the next period continue tracking the code execution dimension, and if it stays below 80 for two consecutive days, launch deep retesting; the current data does not support downgrading the long-term assessment of GPT-o3's code execution ability.
Data source: YZ Index | Run #345 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接