Grok 4 Code Execution Plunges 21.8 Points, Main Leaderboard Falls from 89.15 to 84.99

Grok 4's code execution score in today's Smoke evaluation fell from 94.50 to 72.70, its material constraint score rose from 82.60 to 100.00, and the overall main leaderboard retreated from 89.15 to 84.99.

Score Comparison and Data Facts

Yesterday's code execution score was 94.50 and material constraint was 82.60; today's code execution is 72.70 and material constraint is 100.00. Engineering judgment rose from 75.00 to 86.10, and task expression rose from 86.10 to 91.70. The main leaderboard is composed solely of weighted code execution and material constraint scores, so today's 4.2-point decline stems directly from the 21.8-point plunge in the code execution dimension alone.

Cause Analysis: Question Sampling Variance or Model Degradation

The Smoke evaluation uses only 10 questions per day, with 2 questions per dimension, giving each question extremely high score weight. A single mistake on one code execution question can drag the dimension down by more than 20 points. Today's code execution and material constraint scores moved sharply in opposite directions, indicating that the model's performance differences across question types were amplified rather than reflecting overall capability degradation. The simultaneous rise in engineering judgment and task expression further supports the random sampling variance explanation.

If the model had truly degraded, multiple dimensions would typically decline in tandem, but today's perfect material constraint score shows that the model's ability to handle material fidelity remains unaffected. The data only reflects single-day small-sample fluctuation, with no evidence of consecutive multi-day declines in the same dimension, so systematic degradation cannot be confirmed.

Implications for Users

Development teams that rely heavily on code execution should note that Grok 4's code execution score in the Smoke evaluation can fluctuate by as much as 21.8 points in a single day, meaning output inconsistency may occur in short-cycle tasks. The perfect material constraint score indicates stable performance in scenarios requiring strict adherence to input materials without fabrication, making the model well suited for tasks such as document proofreading and contract review.

The rising engineering judgment and task expression scores suggest the model retains advantages in scenarios requiring engineering decisions and task decomposition. Developers relying on this model can prioritize it for workflows with high material constraint requirements, while assigning high-risk code generation tasks to models with more stable scores.

Strategic Assessment

Based on today's data, the code execution dimension of Grok 4 shows notable fluctuation under a small sample, and the 4.2-point drop on the main leaderboard is worth recording. However, the perfect material constraint score and improvements in secondary metrics indicate no systemic problems in overall model capability. If the code execution score rebounds above 90 in the next Smoke evaluation, this episode can be confirmed as sampling variance; if it remains below 80, further verification of the model's stability on code execution tasks will be required.

Current data does not support any conclusion that overestimates or underestimates Grok 4 overall; it only shows high variance in the daily 10-question quick test. Enterprises selecting models should continue tracking main leaderboard and code execution score changes over three consecutive days to obtain more reliable signals.


Data source: YZ Index | Run #304 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!