Grok 4 Material Constraint Drops 17.6 Points; Main Leaderboard Falls Just 1.8 Points

In today's Smoke evaluation, Grok 4's material constraint score fell from 82.60 to 65.00, while the main leaderboard overall only dropped from 82.99 to 81.23.

Score Comparison Data

Code execution rose from 83.30 to 94.50, engineering judgment climbed from 30.60 to 100.00, task expression fell from 75.00 to 70.00, and the integrity rating changed from pass to warn. The 17.6-point decline in the single dimension of material constraint far exceeds the 1.8-point overall change on the main leaderboard.

Analysis of Fluctuation Causes

The Smoke evaluation consists of only 10 questions per day, with 2 questions covering the material constraint dimension. A score of 65.00 means the average score on those 2 questions was significantly lower than yesterday's. Code execution improved by 11.2 points over the same period, indicating the model performed more consistently on code-related tasks. The drop in material constraint score most likely stems from the 2 material fidelity questions drawn that day being more difficult or having stricter constraints, rather than a regression in the model's overall capability. The sharp rise in engineering judgment also confirms that the model has not shown systemic issues on side-leaderboard tasks.

The integrity rating shifting from pass to warn suggests possible minor factual deviations or constraint violations in this round, which is directly tied to the material constraint score. This dimension primarily assesses how faithfully the model adheres to given materials; a score of 65.00 indicates the model noticeably deviated or over-generated on at least one question.

Specific Impact on Users

Developers who depend on material constraint should add manual verification steps when generating reports, contract summaries, or technical documentation. Teams focused on code execution can continue using Grok 4, whose 94.50 score beats yesterday's performance. In scenarios demanding high material fidelity, such as legal text processing and academic citation generation, the current 65.00 level may cause outputs to drift from source materials, increasing rework risk.

Strategic Assessment

The main leaderboard's modest 1.8-point decline shows overall capability has not collapsed, but the 17.6-point drop in material constraint combined with the integrity rating's warn signal forms a clear indicator that needs to be verified in the next evaluation. If material constraint rebounds above 80 points next time, this is more likely due to question-drawing fluctuations; if it stays below 70 points, a genuine stability decline in the model's material constraint capability should be considered. Current data does not support declaring Grok 4's material constraint capability systemically degraded, but it does constitute an anomaly worthy of close monitoring.

Engineering judgment rising to 100.00 shows the model still has headroom on side-leaderboard tasks, and the modest 5-point decline in task expression has limited impact. Considering the main leaderboard score of 81.23 and material constraint score of 65.00, enterprises should temporarily lower Grok 4's priority in material-sensitive scenarios and wait for subsequent Smoke evaluation data to confirm the trend.


Data Source: YZ Index | Run #266 | View Raw Data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!