Grok 4 Materials Constraint Plunges 22.2 Points, Main Leaderboard Slips 5.2

In today's Smoke evaluation, Grok 4's materials constraint score fell from 100.00 yesterday to 77.80, a drop of 22.2 points, while its main leaderboard score declined from 93.79 to 88.64.

Score Comparison Data

Code execution rose from 88.70 to 97.50, up 8.8 points; engineering judgment fell from 75.00 to 30.60, down 44.4 points; task expression rose from 45.80 to 95.00, up 49.2 points. The integrity rating remained at pass.

Analysis of the Volatility

The Smoke evaluation uses only 2 questions per dimension per day, 10 in total. The simultaneous sharp declines in materials constraint and engineering judgment, alongside equally large gains in code execution and task expression, indicate that the score changes most likely stem from the randomness of question sampling rather than any systematic degradation in model capability. If one of the two materials constraint questions demands strict fidelity to the source text and the model engages in even slight rewriting, that dimension can fall straight from a perfect score to 77.80.

Implications for Users

Scenarios that place heavy emphasis on source fidelity (such as legal contract summarization or verifying academic citations) require an additional round of human review. Developers who rely on code execution can continue to work with today's performance, but should note that the 30.60 score on the engineering judgment side may reflect instability in the model's multi-step reasoning consistency.

Strategic Assessment

Based on a single day's data, a 5.2-point decline on the main leaderboard is not enough to conclude that the model has genuinely degraded. We recommend focusing on whether materials constraint and engineering judgment recover in tandem in the next round; only if both remain below 80 for two consecutive days should a deeper re-test be triggered. The current fluctuation falls within the normal range of question sampling and does not warrant immediately adjusting selection priorities.

Overall, Grok 4's performance today shows the classic signature of high variance in a small sample. The opposing movements of the 77.80 materials constraint score and the 97.50 code execution score further confirm that differences in question content are the main driver. When enterprises select models, the Smoke evaluation should serve as a reference signal rather than the sole basis; the true level of a model should be judged in combination with multi-day averages.


Data source: YZ Index | Run #318 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!