In today's Smoke evaluation, Grok 4's code execution score fell from 95.30 to 64.00, a drop of 31.3 points, while material constraints rose from 48.40 to 83.10, a gain of 34.7 points. Its main leaderboard score slipped marginally from 74.20 to 72.60.
Score Comparison and Key Facts
Engineering judgment rose from 50.00 to 80.60, task expression rose from 75.00 to 81.10, and the integrity rating held at pass. The main leaderboard is a weighted composite of code execution and material constraints; today the two dimensions moved in opposite directions by similar magnitudes, leaving the main leaderboard only slightly lower.
Analysis of the Volatility
The Smoke evaluation uses only 2 questions per dimension per day, so each individual question carries heavy weight. Code execution fell from 95.30 to 64.00, most likely because today's drawn questions placed higher demands on code generation or execution paths, whereas yesterday's questions were relatively easier. Material constraints rose sharply at the same time, indicating that the same set of questions was easier to score on for citation fidelity or constraint compliance. This is not an across-the-board degradation of model capability, but random fluctuation caused by the question draw.
If this were genuine degradation, material constraints and engineering judgment would not normally rise sharply in tandem. In today's data, both material constraints and engineering judgment posted gains of more than 30 points, pointing to a change in question characteristics rather than a systemic problem with model parameters or training data.
What It Means for Users
Teams that rely heavily on code execution should apply additional manual verification to Grok 4 outputs today and in the near term, especially in scenarios involving multi-step calculations or complex scripts. The material constraints score rising to 83.10 indicates the model performs more steadily on tasks requiring strict adherence to input materials or formatting constraints, making it suitable for document generation, compliance checks and similar use cases.
For developers who use both the code execution and material constraints dimensions, today's main leaderboard score of 72.60 remains in a usable range, but volatility in a single dimension may amplify failure rates on specific tasks.
Strategic Assessment
Given today's opposing score movements, the slight decline in the main leaderboard was driven mainly by the single code execution dimension rather than an overall drop in model capability. The simultaneous rise in engineering judgment and task expression further supports the conclusion that the fluctuation stems from the question draw rather than model degradation. The next evaluation should focus on whether code execution recovers to the 90-plus range; if it stays low for two consecutive days, reassessment will be needed.
Current data does not support a long-term downgrade of Grok 4's coding capability; a single-day drop of 31.3 points falls within the normal fluctuation range of the Smoke evaluation. Selection teams are advised to keep tracking the standard deviation of scores on the same dimension over the next 3-5 days rather than adjusting deployment strategy based on a single day's data.
Data source: YZ Index | Run #361 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接