GPT-o3 scored 79.28 points on today's Smoke evaluation main leaderboard, a decrease of 13.9 points from yesterday's 93.16.
Key Data Comparison
The code execution dimension fell from 95.00 to 81.80 points, down 13.2 points; the material constraint dimension fell from 90.90 to 76.20 points, down 14.7 points. The engineering judgment side leaderboard dropped from 94.50 to 63.90 points, down 30.6 points; the task expression side leaderboard rose from 91.70 to 100.00 points, up 8.3 points. The integrity rating remains "pass."
Possible Causes of the Score Decline
The Smoke evaluation includes only 10 questions per day, with 2 questions per dimension, so the small sample size means single-day fluctuations are within the normal range. The simultaneous drops of more than 13 points in both code execution and material constraints most likely stem from the randomly selected questions for the day being more difficult on the corresponding dimensions, or from the model's unstable responses to specific question types. The sharp 30.6-point fluctuation on the engineering judgment side leaderboard further supports that sampling factors are the dominant cause, rather than an overall degradation of model capability. The 8.3-point rise on the task expression side leaderboard also confirms that different dimensions are affected by question selection in different directions.
Specific Implications for Users
Development teams that rely heavily on code execution should note that GPT-o3's single-day score of 81.80 on that dimension may be below its long-term average, and it is advisable to add manual verification steps in production environments. For scenarios sensitive to material constraints, such as document processing tasks requiring high fidelity to long texts, today's 76.20 score suggests the model may produce more deviations from the source. The 63.90 score on the engineering judgment side leaderboard indicates that output consistency may currently be insufficient in scenarios requiring AI-assisted architectural decision-making.
Strategic Assessment
Based on the comparison of current scores, the 13.9-point decline on the main leaderboard is primarily driven by two main dimensions—code execution and material constraints. Combined with the small-sample nature of the Smoke evaluation, this change is more likely due to question sampling fluctuation rather than genuine model degradation. The integrity rating remains "pass," indicating the model has not exhibited systematic integrity issues. It is recommended that the next evaluation cycle focus on whether code execution and material constraints recover to the 90+ range; if they remain low for two consecutive days, further verification is needed.
Overall, GPT-o3's performance today does not provide sufficient evidence to support the conclusion of a permanent decline in model capability, but users with high dependence on code execution and material constraints should increase their monitoring frequency in the short term.
Data source: YZ Index | Run #256 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接