Gemini 3.1 Pro Main Score Plunges 10.8 Points; Material Constraint Drops 14.4 in a Single Day

Gemini 3.1 Pro's main score in today's Smoke evaluation fell from yesterday's 96.60 to 85.83, a drop of 10.8 points.

Breakdown of Core Dimension Changes

The code execution dimension fell from 99.30 to 91.50, a drop of 7.8 points. The material constraint dimension fell from 93.30 to 78.90, a drop of 14.4 points. Engineering judgment (side ranking, AI-assisted evaluation) fell from 75.00 to 50.00, a drop of 25 points. Task expression rose from 85.00 to 91.70, an increase of 6.7 points. Integrity rating remained at pass.

Data Facts and Possible Causes

Smoke evaluation only includes 10 questions per day, 2 questions per dimension, so single-day fluctuations are within the normal range. However, the material constraint dimension dropped 14.4 points in one day, far exceeding the 7.8-point drop in code execution, suggesting that this round's randomly selected questions may have concentrated on exposing the model's weaknesses in material fidelity. Engineering judgment (side ranking, AI-assisted evaluation) dropped 25 points, further indicating that the model made notable errors in scenarios requiring engineering trade-offs.

The task expression dimension actually rose 6.7 points, indicating that the model did not show overall regression in the clarity of task articulation. These two opposing changes point to question-selection volatility rather than genuine model degradation. The simultaneous sharp declines in material constraint and engineering judgment most likely stem from that day's questions including cases that required strict adherence to material boundaries or multi-factor engineering trade-offs, where Gemini 3.1 Pro failed to maintain yesterday's performance.

Implications for Users

For scenarios that heavily depend on material constraints—such as legal contract extraction, product specification verification, and research literature summarization—Gemini 3.1 Pro's performance today signals elevated risk. Developers using this model in pipelines that require strict adherence to source materials should add manual verification steps or switch to models with more stable material constraint scores.

For teams that rely heavily on code execution, the 7.8-point drop is smaller than the material constraint decline, but they should still monitor whether the trend persists over consecutive days. Engineering judgment (side ranking, AI-assisted evaluation) falling to 50 points means that in architecture decision-making and trade-off analysis tasks, the model's output reliability has decreased, and it is recommended to avoid using it directly for production decisions for now.

Strategic Assessment

This round's 10.8-point main score decline was driven primarily by material constraint and engineering judgment, while the rise in task expression suggests that the model's overall capabilities have not systematically declined. In the short term, this can be viewed as question-selection volatility, but the 14.4-point drop in material constraint has reached the threshold requiring tracking.

If the material constraint score fails to recover above 90 in the next round, it may indicate new bottlenecks in Gemini 3.1 Pro's ability to faithfully adhere to source materials. The judgment currently supported by the data is: in scenarios sensitive to material constraints, reliance on Gemini 3.1 Pro should be temporarily reduced, while code execution scenarios can continue to be observed.


Data source: YZ Index | Run #300 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!