In today's Smoke evaluation, Gemini 3.1 Pro's main score dropped from 80.99 to 72.50, a decrease of 8.5 points.
Core Dimension Score Changes
The code execution dimension dropped from 75.00 yesterday to 50.00 today, a decrease of 25 points; the material constraint dimension rose from 88.30 to 100.00, an increase of 11.7 points. Engineering judgment (side ranking, AI-assisted evaluation) dropped from 100.00 to 75.00, and task expression (side ranking, AI-assisted evaluation) dropped from 76.70 to 66.70. The integrity rating remains pass.
Cause Analysis of Fluctuation
The Smoke evaluation only includes 10 questions per day, with 2 questions per dimension, resulting in an extremely small sample size. The single-day 25-point drop in the code execution dimension is most likely due to increased difficulty of sampled questions or inconsistent model outputs in specific code scenarios. The simultaneous rise to full score in the material constraint dimension indicates that the model's performance in citation constraints has not systematically degraded. The opposite changes in the two dimensions point to question sampling fluctuation rather than overall model capability degradation.
Engineering judgment (side ranking, AI-assisted evaluation) and task expression (side ranking, AI-assisted evaluation) both declined, reflecting increased score fluctuation on questions requiring engineering trade-offs and task articulation. The stability dimension formula max(0, 100-stddev×2) has already indicated a large standard deviation in the model's scores on similar questions, and this single-day multi-dimensional fluctuation is consistent with the stability score of 31.7.
Specific Implications for Users
Teams focusing on code execution should add redundant verification when using Gemini 3.1 Pro. A single-day score of 50.00 implies that both code questions may have failed. For scenarios that rely on material fidelity, the model remains usable, as today's 100.00 score shows constraint capability is still at a high level.
For mixed workflows requiring both engineering judgment and task expression, it is recommended to add manual review checkpoints in the production environment to avoid output deviations caused by scores of 75.00 and 66.70.
Strategic Judgment
Based on the current score comparison, Gemini 3.1 Pro's code execution capability is magnified by single-day sampling, and there is insufficient evidence of true degradation. However, if code execution scores remain below 60 for two consecutive days, further verification is needed in the next cycle. The material constraint dimension maintains a full score, indicating it remains a relative strength of the model. The main score of 72.50 is already lower than yesterday; when selecting models, multi-day average scores should be used instead of single-day results.
Data source: YZ Index | Run #232 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接