GPT-o3 Main Score Plummets 13.8 Points, Code Execution Drops from 70.3 to 48.5

GPT-o3's main score in today's Smoke evaluation dropped from 80.61 to 66.86, code execution fell from 70.30 to 48.50, a single-day decline of 21.8 points.

Data Facts

Score comparison shows code execution dropped from 70.30 to 48.50, material constraint fell from 93.20 to 89.30, causing the main score to decline by 13.8 points. Engineering judgment rose from 75.00 to 97.00, task expression dropped from 90.00 to 66.70. Integrity rating remains passing.

Cause Analysis

The Smoke evaluation has only 10 questions per day, with 2 questions per dimension, so question draw fluctuation is the main mechanism. Code execution saw the largest drop, indicating the two questions drawn that day may have exceeded the model's current stable processing range. Material constraint dropped only 3.9 points, suggesting no systemic degradation in the model's basic ability to stay faithful to materials. Engineering judgment surged 22 points, while task expression fell 23.3 points, reflecting uneven difficulty distribution across dimensions rather than a change in overall model capability.

Implications for Users

Teams relying heavily on code execution should note that a score of 48.50 in today's Smoke evaluation means the model may show significant instability when handling complex coding tasks. Material constraint remains at 89.30, so scenarios sensitive to material fidelity can still use the model. Engineering judgment at 97.00 indicates strong performance in engineering decision-making scenarios, but task expression at 66.70 suggests extra validation is needed when generating structured outputs.

Strategic Assessment

This main score decline is primarily driven by single-day fluctuation in code execution, while material constraint changed modestly, so it does not yet constitute a genuine degradation signal. It is recommended to continue monitoring whether code execution recovers to around 70 in the next round. If code execution stays below 50 for two consecutive days, then consider whether the model's capabilities have undergone systemic changes.


Data source: YZ Index | Run #237 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!