Grok 4 Main Ranking Plunges 8.8 Points: Code Execution Drops from 100 to 90.7

Grok 4's main ranking score in today's Smoke evaluation fell from yesterday's 96.99 to 88.14, a single-day decline of 8.8 points. The code execution dimension dropped from 100.00 to 90.70 points, and material constraints fell from 93.30 to 85.00 points, with the two combined dragging down the main ranking.

Score Change Breakdown

The code execution dimension saw the largest decline at 9.3 points. The material constraints dimension also dropped 8.3 points. Engineering judgment rose from 69.50 to 75.00 points, while task expression fell from 95.00 to 91.70 points. The integrity rating remains "pass" and did not trigger any threshold.

Possible Cause Analysis

The Smoke evaluation includes only 10 questions per day, with 2 questions corresponding to each main ranking dimension. Code execution and material constraints each had 2 questions drawn; a single-question mistake can cause 8-10 point fluctuations. Yesterday's full score in code execution versus today's 90.70 most likely stems from harder programming or debugging questions being drawn, rather than a sudden degradation of the model itself. The simultaneous decline in material constraints suggests today's questions may include more tasks requiring strict fidelity to the original text.

Engineering judgment rose in the opposite direction by 5.5 points, indicating that the side-ranking question draw is independent of the main ranking. Task expression dipped slightly by 3.3 points, a smaller decline than the main ranking's core dimensions. Taken together, today's changes are more consistent with question-draw fluctuation than a genuine decline in model capability.

Implications for Users

Development teams that heavily rely on code execution should note that Grok 4's single-run performance on high-difficulty programming tasks may vary considerably. In scenarios sensitive to material constraints—such as legal contract review or technical document extraction—today's 85.00 score suggests adding manual review steps.

The slight uptick in engineering judgment has limited impact on scenarios focused on system design judgment. The overall main ranking of 88.14 remains within the usable range, but consecutive days below 90 will shift selection priorities.

Strategic Assessment

Based on single-day data, the decline in Grok 4's main ranking is primarily driven by the question draw and does not yet constitute a genuine degradation signal. If both code execution and material constraints rebound above 95 in the next Smoke evaluation round, this can be confirmed as normal fluctuation; if they persist below 90, the model's stability on complex constraint tasks will require focused verification.

There is currently no evidence supporting Grok 4 being either overrated or underrated. It is recommended to maintain the observation window for at least 3 consecutive evaluation days.


Data source: YZ Index | Run #300 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!