Gemini 3.1 Pro directly lost 33.5 points in the main rankings of today's Smoke evaluation, with the core reason being a sharp drop in code execution from 100.00 to 20.00. This is not a minor fluctuation, but a near-failure of a core capability in a single day's test.
Topic Sampling or Real Degradation
The Smoke evaluation only includes 10 questions per day, 2 questions per dimension, with a small sample size, making single-day score fluctuations normal. However, the code execution dimension fell by 80 points, far exceeding the normal random range. Materials constraint actually rose from 59.50 to 65.50, indicating no systematic decline in constraint adherence. Engineering judgment increased from 10.00 to 38.40, ruling out the possibility of overall capability collapse.
More notably, the same model achieved a perfect score in code execution yesterday, but failed on both questions today. This points to two possibilities: first, today's selected questions happened to hit the model's current weaknesses; second, after the latest update, the model's robustness in complex code generation and debugging has declined.
Recent Industry Developments as Evidence
Over the past two weeks, Google has made multiple weight adjustments to the Gemini series, focusing on strengthening long-context and multimodal alignment. Historical data shows that such adjustments often come at the cost of code execution capability. A similar situation occurred after the Claude 3.5 Sonnet update in June, when the code dimension also experienced a notable decline for two consecutive weeks.
Based on publicly available model update logs, the most recent weight push for Gemini 3.1 Pro occurred 48 hours ago, prioritizing optimization of mathematical reasoning and safety alignment. Strengthening safety alignment typically increases the model's refusal rate for "high-risk code" requests, which is highly consistent with today's low code execution score.
Should We Be Worried?
Yes. Code execution is one of the only two auditable dimensions in the main rankings, and its weight directly determines the model's usability in engineering scenarios. Although the single-day 80-point drop may be partially attributed to question difficulty, the extreme discrepancy in performance on similar questions over two consecutive days indicates that the model's output consistency has fallen below the pass line.
- If tomorrow's Smoke evaluation code execution score remains below 40, it can be judged as systematic degradation rather than random fluctuation.
- If the score recovers to above 80, this event can be classified as a high-variance event and should not be over-interpreted.
When a model experiences a cliff-like drop of 80 points in a core dimension in a single day, industry analysts should first question not luck, but the inconspicuous "safety alignment" change in the update log.
The only positive signal currently is that the integrity rating has shifted from fail to pass, indicating that the model did not exhibit obvious hallucinations or fabrications in this test. However, this does not mask the substantial decline in code execution capability.
For developers who rely on Gemini for code generation, it is recommended to suspend key task deployments within the next 48 hours, and wait for at least two rounds of Smoke evaluation results before making decisions.
Data source: YZ Index | Run #136 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接