In today's Smoke evaluation, GPT-o3's main leaderboard score fell from 96.01 yesterday to 80.48, a drop of 15.5 points, with the code execution dimension plunging from 97.00 to 64.50, a decline of 32.5 points.
Breaking Down the Data: Changes in Core Dimensions
The main leaderboard consists of only two auditable dimensions: code execution and material constraints. Today, code execution scored 64.50 and material constraints scored 100.00, yielding a simple average of 80.48. Yesterday's corresponding figures were 97.00 for code execution and 94.80 for material constraints, averaging 96.01. Material constraints actually improved by 5.2 points today, indicating no systemic problem with the model's fidelity to source material.
Data from the side leaderboard (AI-assisted evaluation) shows engineering judgment falling from 88.90 to 77.30, while task expression rose from 65.00 to 90.00. The integrity rating remained at pass, with no threshold warnings triggered.
Cause Analysis: Sampling Noise or Real Degradation
The Smoke evaluation uses only 2 questions per dimension per day, 10 questions in total, an extremely small sample. A single-day 32.5-point drop in code execution is most likely due to differences in difficulty or question type introduced by random sampling, rather than parameter-level degradation of the model. The simultaneous 5.2-point rise in material constraints further supports the judgment that this is fluctuation rather than an overall capability decline. The 11.6-point drop in engineering judgment may be linked to the code execution questions, but the 25-point rise in task expression shows the model maintains or improves response quality on expression-type tasks.
If this were genuine degradation, material constraints should not have improved at the same time; the current opposing movement across the two dimensions is more consistent with random fluctuation in a small sample.
Implications for Users
Teams that rely heavily on code execution should add redundant verification steps when deploying GPT-o3, especially in mathematical computation and algorithm implementation scenarios. A material constraints score of 100.00 indicates the model remains suitable for document processing tasks that require strict adherence to source text. An engineering judgment score of 77.30 (side leaderboard, AI-assisted evaluation) suggests that complex system design work requires human review.
Developers who base model selection decisions on a single day's Smoke results may overestimate or underestimate true stability.
Strategic Assessment
Based on today's data, the consistency of GPT-o3's code execution warrants continued tracking in the next period. The main leaderboard score of 80.48 remains above most competitors' baselines, but a 32.5-point swing in a single dimension exceeds the normal range. The coexistence of a perfect material constraints score and a low code execution score shows the model performs unevenly across different capability axes. We recommend that the next Smoke round focus on whether code execution rebounds above 90 points, in order to distinguish incidental fluctuation from systematic change.
There is currently no evidence supporting overall model degradation; users can continue to observe rather than switch immediately.
Data source: YZ Index | Run #335 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接