Grok 4 Leads with 98.35 Points: 2026-07-22 Smoke Quick Test Data Brief

On 2026-07-22, the YZ Index Smoke quick test covered 11 models, with Grok 4 ranking first at 98.35 points. Smoke is a daily 10-question quick test suitable for observing short-term signals and is not equivalent to the Full weekly ranking conclusion.

This Smoke test only covers two main benchmark dimensions: code execution and material constraint. The main benchmark formula is 0.55 × Code Execution + 0.45 × Material Constraint. Due to the small daily sample size, single-day scores are better used as monitoring signals rather than as long-term conclusions about model capabilities.

Daily Ranking

RankModelMain BenchmarkCode ExecutionMaterial ConstraintIntegrity
#1Grok 498.3597100pass
#2Claude Sonnet 4.696.794100warn
#3DeepSeek V4 Pro96.794100pass
#4GPT-5.596.794100pass
#5Claude Opus 4.792.619490.9pass
#6Qwen3 Max92.619490.9pass
#7GPT-o387.9488.387.5warn
#8Doubao Pro87.19775pass
#9GLM-4.6749775fail
#10Gemini 3.1 Pro71.76975pass
#11Gemini 2.5 Pro55.635062.5pass

Data Interpretation

In today's YZ Index Smoke quick test, the top models showed clear differentiation in the combination of code execution and material constraint. Grok 4 led with a main benchmark score of 98.35, code execution 97, material constraint 100, and integrity pass; Claude Sonnet 4.6, DeepSeek V4 Pro, and GPT-5.5 all scored 96.7 on the main benchmark, with code execution 94 and material constraint 100, and integrity scores of warn, pass, and pass respectively. Claude Opus 4.7 and Qwen3 Max scored 92.61 on the main benchmark, with code execution 94 and material constraint 90.9, both integrity pass. Doubao Pro scored 97 in code execution but only 75 in material constraint, with a main benchmark of 87.1; GLM-4.6 also scored 97 in code execution and 75 in material constraint, with a main benchmark of 74 and integrity fail.

Among notable outliers, Gemini 2.5 Pro scored 55.63 on the main benchmark, a drop of 24.6 points compared to the same-methodology run, with code execution 50 and material constraint 62.5; Claude Opus 4.7 rose 18.7 points on the main benchmark, with code execution and material constraint increasing by 19 and 18.3 points respectively; Qwen3 Max rose 13.2 points on the main benchmark, with material constraint up 18.3 points; GLM-4.6 rose 47 points in code execution but its integrity changed from pass to fail; Grok 4 rose 16.7 points in material constraint. GPT-o3 scored 87.94 on the main benchmark, a drop of 8.3 points from the previous test.

The above score changes may result from sampling fluctuations with 10 daily questions, or may reflect real performance differences under specific constraints, requiring confirmation in subsequent same-methodology runs. Smoke quick tests are small-sample single-day signals, and interpretation is limited to day-over-day data comparison, with no long-term extrapolation.

Key Changes

  • Gemini 2.5 Pro: Main benchmark down 24.6 points, code execution -20.8 points, material constraint -29.2 points
  • Claude Opus 4.7: Main benchmark up 18.7 points, code execution +19 points, material constraint +18.3 points
  • Qwen3 Max: Main benchmark up 13.2 points, code execution +9 points, material constraint +18.3 points
  • GLM-4.6: Main benchmark up 11.2 points, code execution +47 points, integrity pass→fail
  • Grok 4: Main benchmark up 9.6 points, material constraint +16.7 points

Signals to Watch

  • GPT-o3: Main benchmark plummeted -8.3 points
  • GLM-4.6: Integrity rating downgraded to Fail (pass→fail)
  • Gemini 2.5 Pro: Main benchmark plummeted -24.6 points

When reading such Smoke briefs, the focus should be on two questions: first, whether a model has shown the same weakness for multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large single-day changes in execution or constraint scores may result from question sampling or may be early signals of real degradation, requiring follow-up runs for confirmation.


Data source: YZ Index | Run #241 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!