Grok 4 Leads with 84.21 Points: 2026-07-24 Smoke Quick Test Data Brief

The 2026-07-24 Winzheng YZ Index Smoke Quick Test covered 10 models, with Grok 4 scoring 84.21 points to top the daily ranking. Smoke is a daily 10-question quick test suitable for observing short-term signals, not equivalent to the full weekly ranking conclusions.

This Smoke evaluation only covers the two main ranking dimensions: code execution and material constraints. The main ranking formula is 0.55 × code execution + 0.45 × material constraints. Due to the small daily sample size, single-day scores are more suitable as monitoring signals rather than long-term conclusions about model capabilities.

Daily Ranking

RankModelMain ScoreCode ExecutionMaterial ConstraintsIntegrity
#1Grok 484.2196.469.3pass
#2Doubao Pro73.927572.6pass
#3GPT-5.573.927572.6pass
#4Claude Opus 4.773.774.672.6pass
#5Gemini 2.5 Pro72.2274.669.3pass
#6DeepSeek V4 Pro68.667560.9pass
#7GPT-o364.927552.6pass
#8Gemini 3.1 Pro59.4758.360.9pass
#9Claude Sonnet 4.658.695069.3pass
#10Qwen3 Max54.915060.9pass

Data Interpretation

Looking at the score structure, Grok 4 achieved a main score of 84.21 with a combination of 96.4 in code execution and 69.3 in material constraints, with its advantage primarily coming from the high code execution score. Doubao Pro and GPT-5.5 both scored 73.92 on the main ranking, with code execution at 75 and material constraints at 72.6, showing a relatively balanced distribution. Claude Opus 4.7 scored 73.7 on the main ranking, with code execution at 74.6 and material constraints at 72.6. Gemini 2.5 Pro scored 72.22 on the main ranking, with code execution at 74.6 and material constraints at 69.3. Among top models, the disparity between code execution and material constraint strengths is evident, with some models complementing their material constraints with stronger code execution performance.

Compared to the previous run under the same methodology, Qwen3 Max dropped 27.1 points in main score, 33.6 points in code execution, and 19.2 points in material constraints; Claude Opus 4.7 dropped 23.3 points in main score, 25.4 points in code execution, and 20.7 points in material constraints; Doubao Pro dropped 14.9 points in main score and 25 points in code execution; Claude Sonnet 4.6 dropped 13.3 points in main score and 25 points in code execution; GPT-o3 dropped 10.1 points in main score and 22.4 points in material constraints. These changes may be due to question sampling fluctuations or incidental variations in single-day signals, requiring confirmation in subsequent runs.

The Smoke Quick Test is a small-sample single-day signal. The above interpretations are based solely on today's data, with language kept restrained and no conclusions drawn about long-term model performance.

Major Changes

  • Qwen3 Max: Main score -27.1, Code execution -33.6, Material constraints -19.2
  • Claude Opus 4.7: Main score -23.3, Code execution -25.4, Material constraints -20.7
  • Doubao Pro: Main score -14.9, Code execution -25
  • Claude Sonnet 4.6: Main score -13.3, Code execution -25
  • GPT-o3: Main score -10.1, Material constraints -22.4

Signals to Watch

  • No release-worthy anomaly signals were retained in this round.

When reading Smoke briefs like this, the focus should be on two questions: First, whether a particular model has repeatedly exhibited the same type of weakness over multiple days; second, whether the integrity rating has shifted from pass to warn or fail. Large daily fluctuations in execution or constraint scores may stem from question sampling or could be early signals of genuine degradation, requiring verification in subsequent runs.


Data Source: YZ Index | Run #244 | View Raw Data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!