GPT-o3 Code Execution Plunges 24.7 Points, Main Leaderboard Falls to 78.13 — Smoke Evaluation Anomaly Warrants Attention

In today's Smoke evaluation, GPT-o3's code execution dimension fell from yesterday's 94.50 to 69.80, a drop of 24.7 points, while the overall main leaderboard declined from 86.04 to 78.13, down 7.9 points.

Data Facts: Clear Divergence Across Dimensions

Material adherence rose from 75.70 to 88.30, up 12.6 points. Engineering judgment fell from 97.00 to 74.80, down 22.2 points. Task expression dropped from 86.70 to 70.80, down 15.9 points. The integrity rating remains pass.

Cause Analysis: Question Sampling Variance More Likely

The Smoke evaluation includes only 10 questions per day, or two questions per dimension. The simultaneous sharp declines in code execution and engineering judgment, set against a rise in material adherence, suggest that the question draw has uneven effects across dimensions. The code execution and engineering judgment questions may have centered on complex computation or multi-step reasoning scenarios, where GPT-o3 made errors on those specific items and produced concentrated score declines. The gain in material adherence suggests that the day's material adherence questions were better matched to the model's current capabilities in either difficulty or type.

If the model were genuinely degrading, multiple core dimensions would typically weaken in tandem, rather than showing a compensatory rise in material adherence. Question sampling variance is therefore the more likely explanation.

Implications for Users

Teams with code-heavy workloads that rely on GPT-o3 for algorithm implementation or data-processing tasks should add manual review checkpoints. The higher material adherence score means GPT-o3 remains usable for document generation or citation scenarios with strict source-fidelity requirements. Given the steep drop in engineering judgment, automated workflows requiring multi-step decisions should receive a lower trust weight in the near term.

Strategic Assessment

Based on a single day of data, the main leaderboard decline was driven mainly by code execution and engineering judgment, while the compensating rise in material adherence indicates no systemic regression in the model's overall capabilities. If code execution and engineering judgment both recover in the next evaluation round, this session can be confirmed as sampling variance; if they stay low, further validation is needed. For now, no downgrade of GPT-o3's overall capability is warranted, but users in code-heavy scenarios should stay alert.

This evaluation round once again shows that Smoke evaluation scores can swing sharply on a single day. The 24.7-point drop in code execution is within an acceptable range given the small sample size, but it is approaching the threshold that warrants close tracking. The 12.6-point rise in material adherence provides countervailing evidence, supporting a variance rather than a degradation interpretation.

For enterprises making model selections, GPT-o3's 88.30 on material adherence remains competitive and keeps it viable as a backup for document-related tasks. Its 69.80 on code execution is already below the everyday level of most mainstream models; developers who depend heavily on code generation should prepare switching plans or run parallel validation.

The synchronized declines in engineering judgment (74.80) and task expression (70.80) further point to concentrated reasoning-chain failures on that day's questions, rather than a problem at the model-parameter level. With the integrity rating holding at pass, basic output-compliance issues are ruled out.

In summary, GPT-o3's anomaly in this Smoke evaluation most likely stems from question sampling variance and does not yet constitute a signal of genuine model degradation. The next round should focus on whether code execution and engineering judgment recover in tandem, so the trend can be confirmed.


Data source: YZ Index | Run #317 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!