Claude Sonnet 4.6 and Grok 4 Tie at 86.25: 2026-09-08 Smoke Quick Test Data Brief

On 2026-09-08, the YZ Index Smoke Quick Test covered 11 models, with Claude Sonnet 4.6 and Grok 4 tied for first at 86.25 points. Smoke is a daily quick test of 10 questions, suitable for observing short-term signals, and is not equivalent to the Full weekly ranking conclusions.

This Smoke evaluation only covers two dimensions of the overall score: code execution and material constraints, with the overall-score formula being 0.55 × Code Execution + 0.45 × Material Constraints. Due to the small daily sample size, single-day scores are more suitable as monitoring signals rather than long-term conclusions about model capabilities.

Daily Ranking

RankModelOverall ScoreCode ExecutionMaterial ConstraintsIntegrity
#1Claude Sonnet 4.686.2575100pass
#2Grok 486.2575100warn
#3Doubao Pro74.3473.875pass
#4DeepSeek V4 Pro72.550100pass
#5Gemini 2.5 Pro70.197564.3pass
#6GLM-4.661.255075warn
#7GPT-o361.255075pass
#8Gemini 3.1 Pro59.0854.864.3pass
#9Claude Opus 4.758.7525100pass
#10GPT-5.556.445064.3pass
#11Qwen3 Max53.0343.864.3pass

Data Interpretation

In today's YZ Index Smoke Quick Test, Claude Sonnet 4.6 and Grok 4 tied for the top overall score at 86.25 points; both scored 75 in code execution and 100 in material constraints, reflecting a combination of a perfect material-constraint score and a moderate code-execution score. Doubao Pro posted an overall score of 74.34, with code execution at 73.8 and material constraints at 75, a relatively balanced structure. DeepSeek V4 Pro posted an overall score of 72.5, with code execution at 50 and material constraints at 100, showing strong material constraints and weaker code execution.

Compared with the previous run under the same criteria, Gemini 3.1 Pro's overall score fell by 24.4 points and code execution fell by 45.2 points; GLM-4.6's overall score fell by 22.2 points, code execution fell by 50 points, but material constraints rose by 11.7 points; Claude Sonnet 4.6's overall score rose by 21 points and material constraints rose by 46.7 points. These changes may stem from question sampling fluctuation, or they may reflect genuine performance differences on specific dimensions, and require confirmation by subsequent runs.

GPT-o3's overall score fell by 16.3 points, code execution fell by 50 points, and material constraints rose by 25 points; GPT-5.5's overall score fell by 15.6 points and code execution fell by 25 points. Smoke is a small-sample, single-day signal; the above changes should not be taken as long-term conclusions.

Key Changes

  • Gemini 3.1 Pro: Overall score down 24.4 points, code execution -45.2 points
  • GLM-4.6: Overall score down 22.2 points, code execution -50 points, material constraints +11.7 points
  • Claude Sonnet 4.6: Overall score up 21 points, material constraints +46.7 points
  • GPT-o3: Overall score down 16.3 points, code execution -50 points, material constraints +25 points
  • GPT-5.5: Overall score down 15.6 points, code execution -25 points

Signals to Watch

  • No publishable anomaly signals were recorded this time.

When reading Smoke briefs like this, the focus should be on two questions: first, whether a particular model has exposed the same type of weakness on multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large single-day swings in execution or constraint scores may stem from question sampling, or they may be early signs of genuine regression, requiring subsequent runs to review.


Data source: YZ Index | Run #314 | View Original Data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!