Gemini 3.1 Pro's material constraint score in today's Smoke evaluation dropped from 90.40 to 72.60, a decrease of 17.8 points, with the main ranking overall falling from 81.93 to 75.90.
Score Changes by Dimension
Code execution rose from 75.00 to 78.60, an increase of 3.6 points. Engineering judgment rose from 88.90 to 100.00, an increase of 11.1 points. Task expression fell from 100.00 to 75.00, a decrease of 25 points. Integrity rating remained at pass.
Analysis of the Material Constraint Drop
The material constraint dimension fell 17.8 points this time, which is the main source of the main ranking decline. The Smoke evaluation draws only 2 questions per day for this dimension, so fluctuations in individual question scores directly amplify the dimension total. The 17.8-point drop corresponds to a significant increase in the average score loss per question in a 2-question test.
Code execution rose 3.6 points in the same period, indicating that the model's performance on execution-type questions did not show a systematic degradation. Engineering judgment reached a perfect score of 100.00, further demonstrating that the model maintains a high level in side-ranking judgment tasks.
Concurrent Decline in Task Expression
The task expression dimension dropped 25 points, a similarly large fluctuation as material constraint. Both side-ranking dimensions experienced notable declines, suggesting that the questions drawn this time may focus on task types that require strict adherence to materials or precise expression.
Determining the Cause of Fluctuation
The Smoke evaluation uses only 10 questions per day, with 2 questions per dimension, so the randomness of drawing leads to a high standard deviation in single-day scores. The 17.8-point drop in material constraint and the 25-point drop in task expression may both stem from the drawn questions having higher demands on material fidelity or expression accuracy, rather than a permanent degradation of model capabilities.
If it were a real degradation, it would typically be accompanied by a concurrent decline in code execution. However, code execution actually rose by 3.6 points this time, so the fluctuation is more likely attributable to question drawing randomness.
Implications for Users
Enterprises heavily reliant on material constraint scenarios, such as contract review, knowledge base Q&A, and long document summarization, should consider temporarily pausing Gemini 3.1 Pro as their preferred model until the next Smoke evaluation results are released.
Developer teams primarily focused on code execution tasks can continue using the model, as the score in this dimension increased rather than decreased, indicating stable execution capability.
Strategic Assessment
This main ranking drop of 6 points is primarily driven by the single dimension of material constraint. Combined with the fact that neither code execution nor engineering judgment deteriorated, Gemini 3.1 Pro's overall capability has not been systematically overestimated or underestimated. If material constraint rebounds above 85 points in the next Smoke evaluation, this can be confirmed as a sampling fluctuation; if it continues to stay below 75 points, multi-day continuous tracking will be required.
The contrast between engineering judgment rising to 100 points and task expression falling to 75 points indicates low correlation between side-ranking dimensions, so each side-ranking score should be evaluated independently during model selection.
Data source: YZ Index | Run #240 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接