Grok 4 Leads with 96.99: 2026-08-29 Smoke Quick Test Data Brief

On 2026-08-29, the YZ Index Smoke quick test covered 11 models, with Grok 4 topping the day at 96.99 points. Smoke is a daily quick test of 10 questions, suitable for observing short-term signals rather than serving as a conclusion equivalent to the Full weekly ranking.

This Smoke evaluation covers only two main board dimensions: code execution and material constraints. The main board formula is 0.55 × Code Execution + 0.45 × Material Constraints. Since the daily sample size is small, single-day scores are better used as monitoring signals rather than definitive long-term judgments of model capability.

Daily Ranking

RankModelMain BoardCode ExecutionMaterial ConstraintsIntegrity
#1Grok 496.9910093.3pass
#2Gemini 3.1 Pro96.699.393.3pass
#3Gemini 2.5 Pro91.9910082.2pass
#4Qwen3 Max91.8290.693.3pass
#5Doubao Pro91.3899.381.7pass
#6Claude Opus 4.783.247593.3pass
#7Claude Sonnet 4.683.247593.3pass
#8GLM-4.683.247593.3pass
#9GPT-o383.247593.3pass
#10DeepSeek V4 Pro78.247582.2pass
#11GPT-5.561.525075.6pass

Data Interpretation

Among the top three on today's main board, Grok 4 achieved a main board score of 96.99 with a combination of 100 for code execution and 93.3 for material constraints, while Gemini 3.1 Pro scored 96.6 with 99.3 for code execution and 93.3 for material constraints. Both show a balanced, high-level pairing of code execution and material constraints. Gemini 2.5 Pro scored 100 for code execution but only 82.2 for material constraints, resulting in a main board score of 91.99, reflecting a structure where code execution stands out while material constraints are relatively weaker. Qwen3 Max scored 90.6 for code execution and 93.3 for material constraints, with a main board score of 91.82, demonstrating a stronger material constraints pairing.

Among notable changes, Gemini 3.1 Pro rose by +33.1 on the main board, +44.8 in code execution, and +18.7 in material constraints; GLM-4.6 rose by +22.2 on the main board, +25 in code execution, and +18.7 in material constraints; DeepSeek V4 Pro rose by +19.8 on the main board, +28 in code execution, and +9.8 in material constraints; and Grok 4 rose by +11.5 on the main board and +20.9 in material constraints. These increases may stem from day-to-day question sampling fluctuations and require follow-up run verification. GPT-5.5, meanwhile, showed a clear decline, dropping -24.4 on the main board and -47 in code execution.

Regarding abnormal signals, Claude Opus 4.7 plunged -10.3 on the main board, Claude Sonnet 4.6 plunged -22 in code execution, GPT-o3 plunged -25 in code execution, and GPT-5.5 plunged -24.4 on the main board. Since the Smoke quick test is a small-sample single-day signal, these changes may be caused by question sampling fluctuations or may reflect genuine regression; all require follow-up run verification to determine stability.

Key Changes

  • Gemini 3.1 Pro: Main board +33.1, Code Execution +44.8, Material Constraints +18.7
  • GPT-5.5: Main board -24.4, Code Execution -47
  • GLM-4.6: Main board +22.2, Code Execution +25, Material Constraints +18.7
  • DeepSeek V4 Pro: Main board +19.8, Code Execution +28, Material Constraints +9.8
  • Grok 4: Main board +11.5, Material Constraints +20.9, Integrity warn → pass

Signals to Watch

  • Claude Opus 4.7: Main board plunged -10.3
  • Claude Sonnet 4.6: Code Execution plunged -22
  • GPT-o3: Code Execution plunged -25
  • GPT-5.5: Main board plunged -24.4

When reading this type of Smoke brief, the focus should be on two questions: first, whether a particular model has exposed the same type of weakness for multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large swings in single-day execution or constraint scores may come from question sampling or may be early signals of genuine regression, requiring follow-up run verification.


Data source: YZ Index | Run #299 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!