GPT-o3 Smoke Evaluation Main Leaderboard Plunges 8.3 Points, Code Execution Drops from 100 to 88.3

In today’s Smoke evaluation, GPT-o3’s main leaderboard score fell from 96.27 to 87.94, a drop of 8.3 points. The code execution dimension fell from 100.00 to 88.30, material constraint from 91.70 to 87.50, engineering judgment from 94.80 to 75.00, task expression rose from 75.00 to 90.00, and integrity rating changed from pass to warn.

Score Change Breakdown

The code execution dimension suffered the largest decline, losing 11.7 points in a single day. Material constraint lost 4.2 points, engineering judgment lost 19.8 points, while task expression instead rose by 15 points. The main leaderboard is composed of a weighted combination of code execution and material constraint, so the overall drop of 8.3 points directly reflects the simultaneous decline in these two dimensions.

Possible Causes Analysis

The Smoke evaluation runs only 10 questions per day, 2 questions per dimension, so question sampling fluctuation is the primary possible factor. Yesterday, both code execution questions scored full marks; today, the scores for the two questions declined, indicating differences in difficulty or scenario coverage. Engineering judgment saw the largest drop, likely because today’s questions emphasized multi-step reasoning or constraint conflicts, exposing instability in the model on that dimension. The slight decline in material constraint suggests some volatility in the model’s fidelity to given material. The rise in task expression indicates that today’s questions were easier for the model to organize and articulate.

Distinguishing real degradation from sampling fluctuation requires multiple days of consecutive data. Currently, only a single-day comparison is available, so it cannot be confirmed whether the model’s capabilities have systematically declined. The integrity rating turning to warn suggests that the model exhibited consistency or compliance issues in some answers, which may be related to the sharp drop in engineering judgment.

Implications for Users

Developer teams heavily reliant on code execution should remain vigilant. Yesterday’s perfect score may have masked potential risks; today’s 88.30 score means that in complex code generation or debugging scenarios, the probability of model errors has increased. The material constraint score of 87.50 places mild pressure on scenarios requiring strict adherence to documents or specifications. The engineering judgment score of 75.00 indicates reduced reliability of model judgment in decision-making tasks that require engineering trade-offs.

The task expression score of 90.00 has relatively little impact on content generation tasks. The integrity rating of warn requires enterprises to add manual review steps before deploying the model in production environments, to avoid boundary violations in model outputs.

Strategic Assessment

Based on the single-day data available, GPT-o3’s main leaderboard decline is primarily driven by fluctuations in code execution and engineering judgment. Question sampling differences are the more likely explanation, but the 19.8-point drop in engineering judgment has already exceeded the normal daily range, warranting continued tracking in the next evaluation. If engineering judgment and code execution both remain at low levels for two consecutive days, it will become necessary to consider the model’s real consistency issues in complex constraint scenarios.

Current data does not support long-term conclusions about the model’s overall capabilities; it only shows that in today’s Smoke evaluation, GPT-o3 exhibited significant fluctuations in the code execution and engineering judgment dimensions. Teams relying on this model should await the next round of evaluation results before deciding whether to adjust their usage strategies.


Data source: YZ Index | Run #241 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!