In today's Smoke evaluation, Grok 4's material constraint score dropped directly from yesterday's 100.00 to 84.20, a decline of 15.8 points, causing its overall main leaderboard score to slip from 95.44 to 90.86.
Core Data Comparison
The code execution dimension rose from 91.70 to 96.30, an increase of 4.6 points; material constraint fell from 100.00 to 84.20, a decline of 15.8 points; engineering judgment (side leaderboard, AI-assisted evaluation) dropped from 79.90 to 61.10, a decline of 18.8 points; task expression (side leaderboard, AI-assisted evaluation) remained unchanged at 70.00; the final main leaderboard result was 90.86; the integrity rating changed from warn to pass.
Analysis of Score Fluctuation Causes
The Smoke evaluation has only 10 questions per day, 2 per dimension, and daily sampling variation itself can amplify score changes. The simultaneous sharp declines in material constraint and engineering judgment suggest that the questions drawn this time may demand higher material fidelity, or that the model showed significant inconsistency in complying with citation constraints. The rise in code execution score indicates that the model's answer stability in that dimension is relatively good, rather than an overall capability regression.
The material constraint dimension directly affects the main leaderboard calculation, and the 15.8-point drop is sufficient to explain the 4.6-point decline in the main leaderboard. Although the 18.8-point drop in engineering judgment is on the side leaderboard, its synchronized change with material constraint suggests that the model may exhibit more deviations or omissions when handling tasks that require strictly following the given materials.
Specific Implications for Users
In scenarios that rely heavily on material constraints, such as contract review, policy interpretation, and long-document summarization, Grok 4's performance today suggests the need for additional human review. Its code execution score of 96.30 remains a useful reference for developers writing and debugging code, but the material constraint level of 84.20 reduces its reliability in tasks requiring strict citation of source text.
The integrity rating changed to pass, indicating that the model showed no obvious violations or hallucinated outputs in this evaluation. This is a positive signal for teams requiring compliant output, but it does not offset the usage risk brought by the decline in material constraints.
Strategic Judgment
Based on single-day data, the main leaderboard decline is mainly determined by material constraints, while the improvement in code execution shows clear differences across dimensions. The simultaneous drop in engineering judgment and material constraints suggests that this question draw may have amplified the model's instability in constraint compliance, rather than indicating sustained degradation.
This fluctuation range has exceeded the common range for Smoke evaluations. It is recommended that the next evaluation focus on whether material constraint rebounds above 95 points. If it remains below 85 points for two consecutive days, Grok 4 should be temporarily removed from the first-choice list for material-sensitive tasks.
Current data does not support interpreting this decline as an overall capability drop. It can only be judged that the material constraint dimension showed significant fluctuation under this draw. For selection teams relying on main leaderboard rankings, Grok 4's main leaderboard score of 90.86 today is already below yesterday's level. When using it, teams need to distinguish between its actual performance in code execution and material constraints according to the scenario.
Data source: YZ Index | Run #334 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接