Grok 4's code execution score in today's Smoke evaluation dropped from 92.00 to 72.50, a decrease of 19.5 points, while material constraint rose from 60.90 to 84.10, an increase of 23.2 points. The main leaderboard score slightly decreased from 78.01 to 77.72.
Direct Data Comparison of Score Changes
Code execution dimension: yesterday 92.00, today 72.50, difference -19.5; material constraint dimension: yesterday 60.90, today 84.10, difference +23.2; engineering judgment dimension: yesterday 94.50, today 75.00, difference -19.5; task expression dimension: yesterday 76.70, today 86.70, difference +10; main leaderboard: yesterday 78.01, today 77.72, difference -0.3. Integrity rating remains pass.
Mechanism Analysis of Fluctuation Causes
The Smoke evaluation uses only 2 questions per dimension per day, totaling 10 questions. With a high weight per question score, the randomness of question sampling directly amplifies the standard deviation. Both code execution and engineering judgment dropped by 19.5 points, suggesting that the sampled questions may have higher requirements for code paths or logical chains. Material constraint rose by 23.2 points, indicating that the other two questions may have placed greater emphasis on citation fidelity or constraint compliance, with the model showing improved performance on these samples. The main leaderboard only dropped by 0.3 points, indicating that the complementary changes between the two core dimensions had a limited weighted impact overall.
Current data cannot distinguish between question sampling fluctuation and genuine model degradation. If large fluctuations in the same direction occur over consecutive days, it would more likely point to capability changes. Single-day data only supports the conclusion that sampling effects are the primary factor.
Scenario Implications for Users
Teams that heavily rely on code execution should conduct additional validation of Grok 4 outputs in today's scoring environment, especially for scenarios involving multi-step calculations or API calls. The material constraint score rising to 84.10 provides higher credibility for document generation tasks that require strict citation of sources or avoidance of hallucinations. Engineering judgment dropping to 75.00 means developers who rely on the model for architectural decisions or trade-off analysis should increase manual review steps.
Strategic Assessment and Validation Signals for the Next Period
The marginal 0.3-point drop in the main leaderboard indicates that Grok 4's overall competitiveness has not been substantially impacted, but the sharp complementary changes in code execution and material constraint warrant continued tracking in the next period. If code execution recovers and material constraint recedes in the next evaluation, this will confirm the current fluctuation as sampling-driven. If code execution consistently remains below 80 points, it will be necessary to assess whether the model is entering a capability adjustment window. Current data supports treating Grok 4's code execution capability as a key monitoring item, rather than downgrading its overall priority.
Data source: YZ Index | Run #255 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接