Claude Sonnet 4.6 Leads with 94.61 Points: 2026-08-20 Smoke Quick Test Data Briefing

On 2026-08-20, the YZ Index Smoke quick test covered 11 models, and Claude Sonnet 4.6 topped the day's rankings with 94.61 points. Smoke is a daily 10-question quick test suitable for observing short-term signals, and is not equivalent to the Full weekly ranking conclusions.

This Smoke evaluation only covers two main leaderboard dimensions: code execution and material constraint. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraint. Due to the small daily sample size, single-day scores are better used as monitoring signals rather than long-term conclusions about model capabilities.

Daily Rankings

RankModelMain LeaderboardCode ExecutionMaterial ConstraintIntegrity
#1Claude Sonnet 4.694.619297.8pass
#2Doubao Pro92.269588.9pass
#3Claude Opus 4.790.169287.9pass
#4Grok 489.8883.497.8pass
#5Qwen3 Max89.2290.387.9pass
#6GPT-5.588.279283.7pass
#7GPT-o384.6376.394.8pass
#8Gemini 3.1 Pro79.516794.8pass
#9Gemini 2.5 Pro62.424583.7pass
#10DeepSeek V4 Pro59.514280.9pass
#11GLM-4.656.164569.8warn

Data Interpretation

In today's YZ Index Smoke quick test, Claude Sonnet 4.6 ranked first on the main leaderboard with 94.61 points, featuring a balanced combination of 92 in code execution and 97.8 in material constraint; Doubao Pro scored 92.26 on the main leaderboard, with 95 in code execution and 88.9 in material constraint, showing a structure stronger in code execution; Claude Opus 4.7 scored 90.16 on the main leaderboard, with 92 in code execution and 87.9 in material constraint, also leaning toward code execution. Grok 4 scored 89.88 on the main leaderboard, with 83.4 in code execution and 97.8 in material constraint, showing relatively stronger material constraint. The score differences among top models under the 0.55×Code Execution + 0.45×Material Constraint formula mainly stem from the varying strength combinations of these two indicators.

Compared with the previous run under the same criteria, Gemini 3.1 Pro's main leaderboard score rose by 31.2 points, with code execution +20 and material constraint +44.8; Doubao Pro's main leaderboard rose by 14.8 points, with code execution -5 and material constraint +38.9; GLM-4.6's main leaderboard rose by 14.9 points, with code execution +20 and material constraint +8.7, and its integrity rating changed from pass to warn. Gemini 2.5 Pro's main leaderboard fell by 11.8 points, with code execution -28.5 and material constraint +8.7; GPT-5.5's main leaderboard fell by 10.1 points, with code execution -5 and material constraint -16.3. These changes may stem from question sampling fluctuations, or may reflect genuine single-day performance, and require follow-up runs for confirmation.

Regarding anomaly signals, Claude Opus 4.7's main leaderboard score plunged by 8.2 points, GPT-5.5's main leaderboard plunged by 10.1 points, GPT-o3's code execution plunged by 22.2 points, and Gemini 2.5 Pro's main leaderboard plunged by 11.8 points. As Smoke is a small-sample single-day signal, the above anomalies are not yet sufficient to distinguish between random fluctuations and model degradation; it is recommended to observe stable trends through multiple rounds of retesting.

Key Changes

  • Gemini 3.1 Pro: Main leaderboard +31.2 points, code execution +20, material constraint +44.8
  • GLM-4.6: Main leaderboard +14.9 points, code execution +20, material constraint +8.7, integrity pass→warn
  • Doubao Pro: Main leaderboard +14.8 points, code execution -5, material constraint +38.9
  • Gemini 2.5 Pro: Main leaderboard -11.8 points, code execution -28.5, material constraint +8.7
  • GPT-5.5: Main leaderboard -10.1 points, code execution -5, material constraint -16.3

Signals to Watch

  • Claude Opus 4.7: Main leaderboard plunged -8.2 points
  • GPT-5.5: Main leaderboard plunged -10.1 points
  • GPT-o3: Code execution plunged -22.2 points
  • Gemini 2.5 Pro: Main leaderboard plunged -11.8 points

When reading this type of Smoke briefing, the focus should be on two questions: first, whether a particular model has exposed the same type of weakness for multiple consecutive days; second, whether the integrity rating has shifted from pass to warn or fail. Large daily fluctuations in execution or constraint scores may come from question sampling, or may be early signals of genuine degradation, requiring subsequent runs for verification.


Data source: YZ Index | Run #286 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!