Gemini 2.5 Pro Code Execution Dropped 24.6 Points in a Single Day; Overall Ranking Slid 6.5 Points

In today's Smoke evaluation, Gemini 2.5 Pro's code execution score dropped from 74.60 to 50.00 points, a decrease of 24.6 points, and its overall ranking also fell from 76.49 to 69.98 points.

Score Change Details

Material constraint score rose from 78.80 to 94.40 points, an increase of 15.6 points. Engineering judgment increased from 75.00 to 88.90 points, up 13.9 points. Task expression slightly dropped from 51.70 to 50.00 points. Integrity rating remained pass.

Analysis of Fluctuation Causes

The Smoke evaluation only uses 2 questions per dimension per day, resulting in a very small sample size. If one of the two questions in the code execution dimension receives no points at all, it can cause a single-day fluctuation of about 25 points. Today's score of 50.00 corresponds exactly to the lowest range where both questions fell short, consistent with an extreme outcome due to random sampling.

At the same time, the material constraint score increased by 15.6 points, indicating the model performed better on the other two questions related to material fidelity. The reverse changes in the two dimensions suggest differences in question content rather than model parameter-level degradation. The concurrent rise of 13.9 points in engineering judgment further supports the interpretation of sampling-induced fluctuation.

Implications for Users

Teams heavily reliant on code execution should add manual verification when using Gemini 2.5 Pro's outputs today, especially in scenarios involving multi-step reasoning or API calls. With a material constraint score of 94.40, the model remains highly usable for tasks that require strict adherence to input documents.

When selecting models, developers should treat the code execution score of 50.00 as a single-day extreme value rather than a new baseline. Only by observing the standard deviation of scores for the same dimension over multiple consecutive days can true stability be assessed.

Strategic Assessment

Based on available comparative data, Gemini 2.5 Pro's code execution capability is more likely experiencing a temporary dip due to question sampling rather than actual model degradation. Although the overall ranking of 69.98 is lower than yesterday's, the simultaneous improvement in material constraint and engineering judgment indicates that the overall capability distribution remains within the normal range.

In the next evaluation, it is crucial to verify whether code execution rebounds above 70 points. If it stays around 50 points for two consecutive days, further investigation is needed to rule out systematic issues. For now, the signals only warrant attention and do not constitute a reason to downgrade long-term assessment.


Data source: YZ Index | Run #238 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!