In today's Smoke evaluation, Grok 4's main leaderboard score fell from 90.86 yesterday to 80.89, with the code execution dimension plunging from 96.30 to 69.50, while the material constraints dimension rose from 84.20 to 94.80.
Breakdown of Data Facts
The main leaderboard consists of only two dimensions: code execution and material constraints. Today, code execution lost 26.8 points in a single day, directly pulling the main leaderboard down by about 13.4 points; material constraints rose 10.6 points over the same period, partially offsetting the decline, for a final net drop of 10 points on the main leaderboard. The Engineering Judgment side leaderboard (AI-assisted evaluation) fell from 61.10 to 36.10, while the Task Expression side leaderboard (AI-assisted evaluation) rose from 70.00 to 90.00. The integrity rating remained pass.
Smoke evaluation includes only 2 questions per dimension each day, 10 questions in total. Each question carries high score weight; differences in question sampling can cause fluctuations of more than 20 points.
Cause Analysis
The decline in the code execution dimension far exceeds the rise in material constraints, most likely because the two code questions drawn today matched the model's current weaknesses in difficulty or type. The rise in the material constraints score indicates that, within the same set of questions, the model actually performed better on requirements for fidelity to materials, ruling out systemic degradation of overall capabilities.
The simultaneous 25-point and 20-point fluctuations on the Engineering Judgment and Task Expression side leaderboards further confirm that this evaluation was significantly affected by question randomness, rather than changes in model parameters or training.
Implications for Users
Teams that rely heavily on code generation should note: Grok 4's code execution score in the Smoke evaluation has shown a single-day fluctuation of 26.8 points, meaning unpredictable failures may occur on certain programming tasks. The material constraints score rebounded to 94.80, showing that it remains highly reliable in scenarios requiring strict adherence to given documents or instructions.
For developers using multiple models at the same time, Grok 4's performance today suggests it can serve as an option for material processing, but the code execution step still requires human review or parallel calls to other models.
Strategic Assessment
Based on the current score comparison, Grok 4's main leaderboard decline this time was mainly driven by single-day sampling fluctuation in the code execution dimension, rather than genuine model degradation. The simultaneous improvement in the material constraints dimension also supports this judgment. It is recommended that the next Smoke evaluation continue to track code execution scores; only if the score is below 80 for two consecutive days should a change in model capability be considered.
Current data does not support a conclusion that Grok 4's coding ability has declined; the inherent fluctuation of a single-day 10-question test is a more reasonable explanation.
Data from: YZ Index | Run #335 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接