Gemini 2.5 Pro scored 79.41 on the main leaderboard in today's Smoke evaluation, down from 88.53 yesterday, a drop of 9.1 points.
Core Data Breakdown
The main leaderboard is composed solely of two weighted dimensions: code execution and material constraint. Today's code execution score was 70.00, down 29.6 points from yesterday's 99.60; the material constraint score was 90.90, up 15.9 points from yesterday's 75.00. With the two dimensions moving in opposite directions, the net effect dragged the main leaderboard down 9.1 points. Engineering judgment and task expression belong to the side leaderboard (AI-assisted evaluation), rising to 100.00 and 65.00 respectively today, and are not counted in the main leaderboard ranking.
Cause Analysis: Sampling Fluctuation or Genuine Degradation
The Smoke evaluation includes only 10 questions per day, with 2 questions per dimension. The code execution score fell from 99.60 to 70.00 in a single day, meaning at least 1 of the 2 questions involved a clear error or failed to meet the full-score standard. Under small-sample sampling, such a magnitude can be entirely attributed to the random distribution of question difficulty rather than structural degradation of model capability. The material constraint score rose 15.9 points in tandem, also confirming that the day's question sampling had opposite effects on different dimensions. The integrity rating remains "pass," with no answer-consistency issues observed.
What This Means for Users
Teams that rely heavily on code execution should take note: Gemini 2.5 Pro scored only 70.00 on today's 2 code execution questions, indicating a relatively high failure probability in algorithm implementation, debugging, or multi-step reasoning scenarios. The improvement in the material constraint score to 90.90 shows that the model has improved in citing sources and avoiding hallucinations, making it suitable for document generation tasks that demand high output fidelity. The side-leaderboard engineering judgment score of 100.00 indicates the model maintains a high standard on engineering decision-making problems, but the task expression score of 65.00 remains in a low range, suggesting room for improvement in complex instruction-following capability.
Strategic Assessment
Based on today's data, the main leaderboard decline was driven primarily by extreme volatility in the code execution dimension alone, rather than a synchronized drop across all dimensions. With only 2 questions per day, the sample size is too small for a single-day decline of 29.6 points to determine genuine model degradation. However, if the same dimension scores below 85 for two consecutive days, it warrants elevated attention. It is recommended that the next round focus on verifying the distribution of question types in the code execution dimension to distinguish random fluctuation from capability changes.
Data source: YZ Index | Run #277 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接