Gemini 2.5 Pro in today's Smoke evaluation saw its Material Constraints score fall from yesterday's 100.00 to 72.00, a drop of 28 points, while Code Execution rose from 71.90 to 100.00, and the Main Leaderboard score rose from 84.55 to 87.40.
Direct Facts of the Score Changes
This evaluation covered only 10 daily questions, with 2 questions each for Code Execution and Material Constraints. Material Constraints lost 28 points in a single day, Engineering Judgment remained unchanged at 100.00, Task Expression rose from 81.70 to 91.70, and the Integrity Rating remained pass. The Main Leaderboard is weighted from the two dimensions of Code Execution and Material Constraints, with a net increase of 2.9 points today.
Question Draw Volatility or Model Degradation
Smoke evaluation questions are randomly drawn each day, and at a scale of 2 questions per dimension, the score weight of a single question is extremely high. Material Constraints scored full marks yesterday and 72 today, and the difference exactly corresponds to losing points on one question; Code Execution scored 71.90 yesterday and full marks today, likewise consistent with a single question lifting the score. Engineering Judgment and Task Expression did not decline in tandem, indicating that the model's overall output framework has not collapsed. Only the Material Constraints dimension showed -28, pointing to a random strengthening of the questions' demands for material fidelity, rather than parameter-level model degradation.
The sharp opposite swings in Material Constraints and Code Execution are a typical sampling effect in a 10-question rapid evaluation.
Concrete Impact on Users
Scenarios that rely on Material Constraints include long-document summarization, contract clause checking, and citation tracing tasks. Gemini 2.5 Pro's 72-point performance today means that in these scenarios the probability of material omissions or rewrites has increased. Teams that prioritize Code Execution can directly benefit from the 100-point performance, with higher usability that day for algorithm debugging and data processing tasks. Enterprises with both types of scenarios need to prepare dual validation procedures for the same model.
Strategic Judgment
The Main Leaderboard score remains in the 87.40 range, indicating that a 28-point swing in a single dimension has not changed the overall ranking position. Observing the same model's standard deviation in Material Constraints over multiple consecutive days can distinguish sampling noise from true capability drift. Currently, only single-day data supports the conclusion that "random draw caused it," and it has not reached the threshold requiring a downgrade of long-term selection priority. If Material Constraints remain below 85 in the next Smoke evaluation, targeted long-text evidence-alignment testing should be initiated.
Engineering Judgment and Task Expression, the two side leaderboard dimensions, remained stable or improved, further confirming that the main cause is the random difficulty of the Material Constraints questions rather than a decline in the model's general capabilities. When selecting models, enterprises can treat Gemini 2.5 Pro's Material Constraints performance as a high-variance signal and increase the proportion of manual spot checks in key deployment scenarios.
Data source: YZ Index | Run #358 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接