Gemini 2.5 Pro Drops 9.1 Points on Main Leaderboard; Code Execution Plunges 29.6 in Single Day

Gemini 2.5 Pro scored 79.41 on the main leaderboard in today's Smoke evaluation, down from 88.53 yesterday, a drop of 9.1 points.

Core Data Breakdown

The main leaderboard is composed solely of two weighted dimensions: code execution and material constraint. Today's code execution score was 70.00, down 29.6 points from yesterday's 99.60; the material constraint score was 90.90, up 15.9 points from yesterday's 75.00. With the two dimensions moving in opposite directions, the net effect dragged the main leaderboard down 9.1 points. Engineering judgment and task expression belong to the side leaderboard (AI-assisted evaluation), rising to 100.00 and 65.00 respectively today, and are not counted in the main leaderboard ranking.

Cause Analysis: Sampling Fluctuation or Genuine Degradation

The Smoke evaluation includes only 10 questions per day, with 2 questions per dimension. The code execution score fell from 99.60 to 70.00 in a single day, meaning at least 1 of the 2 questions involved a clear error or failed to meet the full-score standard. Under small-sample sampling, such a magnitude can be entirely attributed to the random distribution of question difficulty rather than structural degradation of model capability. The material constraint score rose 15.9 points in tandem, also confirming that the day's question sampling had opposite effects on different dimensions. The integrity rating remains "pass," with no answer-consistency issues observed.

What This Means for Users

Teams that rely heavily on code execution should take note: Gemini 2.5 Pro scored only 70.00 on today's 2 code execution questions, indicating a relatively high failure probability in algorithm implementation, debugging, or multi-step reasoning scenarios. The improvement in the material constraint score to 90.90 shows that the model has improved in citing sources and avoiding hallucinations, making it suitable for document generation tasks that demand high output fidelity. The side-leaderboard engineering judgment score of 100.00 indicates the model maintains a high standard on engineering decision-making problems, but the task expression score of 65.00 remains in a low range, suggesting room for improvement in complex instruction-following capability.

Strategic Assessment

Based on today's data, the main leaderboard decline was driven primarily by extreme volatility in the code execution dimension alone, rather than a synchronized drop across all dimensions. With only 2 questions per day, the sample size is too small for a single-day decline of 29.6 points to determine genuine model degradation. However, if the same dimension scores below 85 for two consecutive days, it warrants elevated attention. It is recommended that the next round focus on verifying the distribution of question types in the code execution dimension to distinguish random fluctuation from capability changes.


Data source: YZ Index | Run #277 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!