Grok 4's main leaderboard score in today's Smoke Evaluation dropped from 89.30 to 78.01, a decline of 11.3 points.
Score Breakdown and Key Changes
The Code Execution dimension fell from 97.80 to 92.00, a drop of 5.8 points; the Material Constraint dimension dropped from 78.90 to 60.90, a decline of 18 points; the Engineering Judgment dimension rose from 81.90 to 94.50, an increase of 12.6 points; and the Task Expression dimension increased from 75.80 to 76.70, a rise of 0.9 points. The main leaderboard is weighted solely by Code Execution and Material Constraint, so a single-day drop of 18 points in Material Constraint directly pulled the main score down by 11.3 points.
Analysis of Volatility Causes
The Smoke Evaluation has only 10 questions per day, with 2 questions corresponding to one main leaderboard dimension, leading to a naturally high standard deviation in daily scores due to the small sample size. The concentrated score loss in the Material Constraint dimension this time is suspected to be related to significant deviations in instruction following or output boundary control on the two constraint-related questions drawn that day. The 5.8-point drop in the Code Execution dimension could similarly be due to random selection of question difficulty or test cases, rather than model ability degradation. The reverse increase of 12.6 points in Engineering Judgment further supports that this change is more likely a fluctuation from question sampling than an actual model degradation.
The single-day 18-point drop in the Material Constraint dimension is the direct cause of this main leaderboard decline.
Implications for Users
For scenarios heavily reliant on Material Constraint, such as document generation that requires strict adherence to format, length, or prohibited items, Grok 4's performance today shows significant consistency issues. In Code Execution scenarios, a score of 92.00 remains in the high range, but it has declined from yesterday's 97.80. Teams relying on precise code generation should increase manual review measures. Engineering Judgment rose to 94.50, indicating that Grok 4 performed better today than yesterday in scenarios requiring engineering trade-offs.
Strategic Assessment
Based on this score comparison, the decline in Grok 4's main leaderboard is primarily driven by a sharp single-day drop in the Material Constraint dimension, with a relatively modest decline in Code Execution. The significant reverse increase in Engineering Judgment suggests notable inconsistency across different dimensions. Current data does not support the conclusion of a systematic degradation in model capability; it is more likely a normal fluctuation due to the small sample sampling in the Smoke Evaluation. If Material Constraint continues to score below 70 in the next Smoke Evaluation, attention priority should be increased; if it recovers to above 75, this decline can be considered an anomaly from a single sampling event.
Integrity rating remains pass, with no threshold issues triggered. The stability dimension did not provide new data this time, so changes in response consistency cannot be directly assessed. Enterprises conducting model selection should continue tracking the Material Constraint dimension's performance over three consecutive days to distinguish random fluctuations from genuine capability changes.
Data Source: YZ Index | Run #254 | View Raw Data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接