GPT-o3's main leaderboard score in today's Smoke evaluation fell from yesterday's 100.00 to 89.92, a decline of 10.1 points, primarily caused by the material constraint dimension dropping from 100.00 to 77.60.
Score Details and Direct Comparison
The code execution dimension scored 100.00 on both yesterday and today, with no change. The material constraint dimension fell from 100.00 yesterday to 77.60 today, a decline of 22.4 points. The engineering judgment dimension rose from 75.00 to 95.80, an increase of 20.8 points. The task expression dimension fell from 90.00 to 85.00, a decline of 5 points. The integrity rating remains pass.
Cause Analysis: Draw Fluctuation or True Degradation
The Smoke evaluation covers only 2 questions per dimension per day. Today's 2 questions in the material constraint dimension directly dragged the dimension down by 22.4 points. The 2 questions in the code execution dimension still earned full marks, indicating that the model shows no systemic issues on code-related tasks. The engineering judgment dimension instead improved by 20.8 points, suggesting that the model's performance on side-leaderboard tasks has gotten better. Task expression dipped slightly by 5 points, a limited decline. Taken together, the material constraint's single-day plunge is most plausibly explained by today's randomly drawn 2 questions placing higher demands on material fidelity, with the model losing points on citation or constraint adherence—rather than an overall capability degradation.
Specific Implications for Users
Teams that depend heavily on code execution can continue to rely on GPT-o3, as this dimension has earned full marks for two consecutive days and execution stability remains unaffected. For scenarios sensitive to material fidelity (e.g., contract extraction, literature summarization, and data citation tasks), today's 77.60 serves as a reminder to add manual verification steps to prevent model output from deviating from the given materials. The engineering judgment side leaderboard rose to 95.80, meaning today's performance is better than yesterday's for developers who need model-assisted solution evaluation.
Strategic Assessment
The 10.1-point drop on the main leaderboard was entirely driven by the 22.4-point decline in the material constraint dimension alone, with no simultaneous deterioration in other core dimensions, so there is no need to immediately conclude that the model has degraded overall. If the material constraint dimension continues to produce similarly low scores in the next Smoke evaluation, focused verification will be required; if it recovers to above 95, today's result can be attributed to draw fluctuation. The conclusion supported by current data is that the material constraint dimension has become the short-term focus of observation for GPT-o3, while the code execution dimension retains its full-mark advantage.
Data source: YZ Index | Run #298 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接