Grok 4 Material Constraint Score Plunges 17.8 Points, Code Execution Rises 18.5 Points, Smoke Benchmark Main Score Still Up 2.2

Grok 4 scored 88.42 points on today's Smoke benchmark main board, up 2.2 points from yesterday's 86.25, but the material constraint dimension fell from 100.00 to 82.20 points, a drop of 17.8 points.

Core Data Comparison

The code execution dimension rose from 75.00 to 93.50 points, an increase of 18.5 points. The material constraint dimension fell from 100.00 to 82.20 points, a decrease of 17.8 points. Engineering judgment fell from 88.90 to 86.10 points, a decrease of 2.8 points. Task expression rose from 86.70 to 90.80 points, an increase of 4.1 points. Integrity rating remains warn.

Direct Facts of Dimension Score Changes

The Smoke benchmark covers only 2 questions per dimension per day, and the small sample size makes single-day fluctuations normal. Material constraint scored a perfect 100.00 yesterday and fell to 82.20 today, indicating that at least one of the day's two questions involved material deviation or constraint violation. Code execution scored 75.00 yesterday and 93.50 today, indicating a notable improvement in code accuracy or execution pass rate on that day's two questions.

Material constraint at 82.20 and code execution at 93.50 appearing on the same day constitute the most striking internal contrast of this evaluation run.

Potential Cause Analysis

Question draw fluctuation is the primary explanation. With two questions randomly selected each day, yesterday's material constraint set featured two fully compliant questions, while today's set included one or two questions requiring strict fidelity to source materials, pulling the score down. For code execution, yesterday's draw included harder or edge-case questions, whereas today's were routine cases, lifting the execution pass rate.

Actual model degradation is less likely. If a single-day drop of 17.8 points in material constraint reflected systematic degradation, other dimensions would typically decline in tandem. Yet engineering judgment fell only 2.8 points, task expression actually rose 4.1 points, and the main board still gained 2.2 points overall—there is no consistent degradation signal.

Concrete Implications for Users

For scenarios that demand high material fidelity—such as legal contract generation, product specification extraction, and research literature summarization—Grok 4's performance today indicates a risk of material deviation in individual responses. Developers should incorporate manual verification steps or multi-turn prompt constraints.

For code-execution-heavy scenarios—such as script generation, data processing pipelines, and algorithm verification—Grok 4's 93.50 today offers a strong pass rate and can be used directly in low-risk tasks.

Strategic Assessment

This 17.8-point swing in material constraint is more likely attributable to variation in the two-question draw than to capability degradation. The main board still rising 2.2 points supports this reading. The next evaluation round needs to examine whether material constraint recovers to above 90 points; if it stays below 85 points for two consecutive days, further consistency verification would be warranted.

The minor movements in engineering judgment and task expression did not weigh on the main board, indicating no systematic shift in core capabilities. Enterprises conducting vendor selection can continue to treat Grok 4 as a candidate for code-execution-priority scenarios, but should add extra monitoring metrics for material-constraint-sensitive use cases.

The daily sample-size limit of the Smoke benchmark means that single-day swings on the order of 17.8 points fall within a normal range. Current data does not support a "model degradation" conclusion; it only supports the observation that "today's question difficulty distribution skewed toward material constraints."


Data source: YZ Index | Run #315 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!