In today's Smoke evaluation, GPT-5.5's code execution score dropped from 94.50 to 69.50, material constraint rose from 76.00 to 100.00, and the main index fell from 86.18 to 83.23.
Data Facts: A Slight Decline After Hedging on the Main Index
The main index comprises only two auditable dimensions: code execution and material constraint. Today, code execution dropped 25 points while material constraint rose 24 points, leaving a net decline of 3 points on the main index after the two offset each other. Engineering judgment held steady at 100.00, while task expression fell from 100.00 to 81.70. The integrity rating remained pass.
The Smoke evaluation includes only 10 questions per day, with 2 per dimension. The daily sample size is extremely small, and a single question error can trigger fluctuations of 20-30 points. The simultaneous large reverse movements in code execution and material constraint are consistent with sampling fluctuation rather than systematic one-directional degradation.
Cause Analysis: Question Sampling Fluctuation Is the Most Likely Explanation
The concentrated point losses in the code execution dimension today likely stem from at least 1 of the 2 questions involving complex multi-step computation or long-chain dependencies. The material constraint dimension simultaneously reached a perfect score, indicating that the model adhered to instruction boundaries more strictly within the same batch of questions. The simultaneous extreme reverse changes in two capabilities within the same model point to differences in question content rather than overall capability degradation caused by model parameters or post-training updates.
The 18.3-point drop in the task expression dimension belongs to the side index (AI-assisted evaluation) and does not enter the main index calculation. The direction of this dimension's fluctuation aligns with the decline in code execution, but the magnitude is smaller, suggesting possible overlapping effects from cross-dimension question overlap.
Implications for Users
Teams that heavily depend on code execution should pay extra attention to single-day fluctuation risk when selecting models. The Smoke evaluation includes only 2 questions per dimension, so a score of 69.50 may simply reflect the difficulty of the questions drawn that day rather than the true lower bound of the model's coding capability. Scenarios that rely on material constraint—such as long instruction following or strictly formatted output—performed better today and can serve as a complementary reference.
For developers using both the code execution and material constraint dimensions, today's data shows a degree of complementarity between the two. The main index fell only 3 points net, indicating that the impact of a single dimension's sharp drop on overall usability has been partially offset.
Strategic Assessment
Based on today's score comparison, the plunge in code execution is more likely attributable to question sampling fluctuation than genuine model degradation. The slight decline in the main index is consistent with the pattern of reverse movements across the two dimensions, and there is currently no evidence supporting sustained capability decline. If code execution rebounds above 90 in the next Smoke evaluation, the fluctuation nature would be further confirmed; if it persistently falls below 75, a larger-sample retest should be initiated.
The engineering judgment maintaining a perfect score indicates that the model's decision consistency at the side index (AI-assisted evaluation) level remains unaffected. The pass integrity rating meets the entry threshold. Stability and usability are operational signals, and since no relevant data was provided in this round, they are not included in the assessment.
Overall, GPT-5.5's performance today constitutes an extreme sample within the normal range of Smoke evaluations. Selection teams can continue to observe data from the next 3-5 days before deciding whether to adjust workflows related to code execution.
Data source: YZ Index | Run #304 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接