GPT-o3 Smoke Evaluation Main Leaderboard Plunges 10.1 Points; Material Constraint Drops 22.4 Points in a Single Day

GPT-o3's main leaderboard score in today's Smoke evaluation fell from yesterday's 100.00 to 89.92, a decline of 10.1 points, primarily caused by the material constraint dimension dropping from 100.00 to 77.60.

Score Details and Direct Comparison

The code execution dimension scored 100.00 on both yesterday and today, with no change. The material constraint dimension fell from 100.00 yesterday to 77.60 today, a decline of 22.4 points. The engineering judgment dimension rose from 75.00 to 95.80, an increase of 20.8 points. The task expression dimension fell from 90.00 to 85.00, a decline of 5 points. The integrity rating remains pass.

Cause Analysis: Draw Fluctuation or True Degradation

The Smoke evaluation covers only 2 questions per dimension per day. Today's 2 questions in the material constraint dimension directly dragged the dimension down by 22.4 points. The 2 questions in the code execution dimension still earned full marks, indicating that the model shows no systemic issues on code-related tasks. The engineering judgment dimension instead improved by 20.8 points, suggesting that the model's performance on side-leaderboard tasks has gotten better. Task expression dipped slightly by 5 points, a limited decline. Taken together, the material constraint's single-day plunge is most plausibly explained by today's randomly drawn 2 questions placing higher demands on material fidelity, with the model losing points on citation or constraint adherence—rather than an overall capability degradation.

Specific Implications for Users

Teams that depend heavily on code execution can continue to rely on GPT-o3, as this dimension has earned full marks for two consecutive days and execution stability remains unaffected. For scenarios sensitive to material fidelity (e.g., contract extraction, literature summarization, and data citation tasks), today's 77.60 serves as a reminder to add manual verification steps to prevent model output from deviating from the given materials. The engineering judgment side leaderboard rose to 95.80, meaning today's performance is better than yesterday's for developers who need model-assisted solution evaluation.

Strategic Assessment

The 10.1-point drop on the main leaderboard was entirely driven by the 22.4-point decline in the material constraint dimension alone, with no simultaneous deterioration in other core dimensions, so there is no need to immediately conclude that the model has degraded overall. If the material constraint dimension continues to produce similarly low scores in the next Smoke evaluation, focused verification will be required; if it recovers to above 95, today's result can be attributed to draw fluctuation. The conclusion supported by current data is that the material constraint dimension has become the short-term focus of observation for GPT-o3, while the code execution dimension retains its full-mark advantage.


Data source: YZ Index | Run #298 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!