Grok 4's code execution score in today's Smoke evaluation fell from 96.70 to 74.30, directly pulling the main leaderboard down from 88.33 to 75.38, a drop of 13 points.
Score Details and Dimension Breakdown
The main leaderboard is weighted from two components: code execution and material constraints. Today's code execution score is 74.30, compared with 96.70 yesterday, a drop of 22.4 points; material constraints fell from 78.10 to 76.70, down only 1.4 points. The Engineering Judgment side leaderboard (AI-assisted evaluation) fell from 72.20 to 36.10, while the Task Expression side leaderboard (AI-assisted evaluation) dropped from 75.00 to 70.80. The integrity rating rose from warn to pass.
Causes of the Fluctuation
The Smoke evaluation has only 2 questions per dimension each day, so each question carries high weight. The code execution dimension had the largest decline, indicating that the programming questions selected today were significantly harder than yesterday's. Material constraints remained relatively stable, indicating that the model's fidelity to provided materials did not show systematic degradation. The sharp drop on the Engineering Judgment side leaderboard may likewise stem from the day's questions placing higher demands on the reasoning chain, rather than from a change in the model's overall capability.
The distinction between genuine degradation and sampling fluctuation lies in cross-dimensional consistency. Material constraints declined only slightly, and the integrity rating actually improved, so there is a lack of evidence for a systematic decline in model capability. The standard deviation of a single-day 10-question test is naturally large; under a small sample, a 22.4-point drop is within the normal range.
Specific Implications for Users
Teams that rely heavily on code generation should add a manual review step under today's conditions, especially for tasks involving complex algorithms or multi-step debugging. Material constraints remain at 76.70, so the impact is limited for document-processing scenarios that require strict quotation of source text. The Engineering Judgment side leaderboard score of 36.10 suggests that, in architecture design tasks requiring long-chain reasoning, current output consistency is already below yesterday's level.
Strategic Assessment
This decline was mainly driven by single-day sampling fluctuation in the code execution dimension; material constraints and the integrity rating did not deteriorate in tandem, so it does not constitute a clear signal of genuine model degradation. If code execution remains below 80 in the next Smoke evaluation, its priority for attention should be raised; if it rebounds above 90, today's drop can be confirmed as an isolated fluctuation.
Data source: YZ Index | Run #330 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接