Claude Opus 4.7 Leads with 95.08 Points: 2026-08-21 Smoke Quick Test Data Brief

On 2026-08-21, the YZ Index Smoke quick test covered 10 models, with Claude Opus 4.7 ranking first for the day at 95.08 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals and not equivalent to the Full weekly ranking conclusion.

This Smoke evaluation only covers two main leaderboard dimensions: code execution and material constraints. The leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraints. Due to the small daily sample size, single-day scores are better used as monitoring signals rather than as long-term conclusions about model capabilities.

Daily Ranking

RankModelMain LeaderboardCode ExecutionMaterial ConstraintsIntegrity
#1Claude Opus 4.795.0898.590.9pass
#2Qwen3 Max79.647585.3warn
#3Doubao Pro78.67583pass
#4Grok 476.3671.981.8pass
#5Claude Sonnet 4.672.1576.666.7pass
#6GPT-5.570.0958.384.5pass
#7DeepSeek V4 Pro68.8858.381.8pass
#8GPT-o362.4756.869.4warn
#9Gemini 3.1 Pro505050warn
#10Gemini 2.5 Pro18.3723.512.1fail

Data Interpretation

Looking at the score structure, Claude Opus 4.7 posted 95.08 on the main leaderboard, with code execution at 98.5 and material constraints at 90.9, forming a relatively balanced profile and taking first place. Qwen3 Max posted 79.64 on the main leaderboard, with code execution at 75 and material constraints at 85.3, showing comparatively stronger material constraint performance. Doubao Pro posted 78.6 on the main leaderboard, with code execution at 75 and material constraints at 83, a similar structure. Grok 4 posted 76.36 on the main leaderboard, with code execution at 71.9 and material constraints at 81.8, also leaning toward material constraints.

Among the models with notable fluctuations, Gemini 2.5 Pro fell 44.1 points on the main leaderboard, 21.5 points on code execution, and 71.6 points on material constraints, with integrity dropping from pass to fail. Gemini 3.1 Pro fell 29.5 points on the main leaderboard, 17 points on code execution, and 44.8 points on material constraints, with integrity turning to warn. Claude Sonnet 4.6 fell 22.5 points on the main leaderboard, 15.4 points on code execution, and 31.1 points on material constraints. GPT-o3 fell 22.2 points on the main leaderboard, 19.5 points on code execution, and 25.4 points on material constraints, with integrity turning to warn. GPT-5.5 fell 18.2 points on the main leaderboard and 33.7 points on code execution. These changes may stem from question sampling fluctuation or may reflect genuine degradation, and require confirmation in subsequent runs.

The Smoke test is a small-sample single-day signal; the interpretation above remains measured and draws no long-term conclusions.

Major Changes

  • Gemini 2.5 Pro: main leaderboard -44.1 points, code execution -21.5 points, material constraints -71.6 points, integrity pass→fail
  • Gemini 3.1 Pro: main leaderboard -29.5 points, code execution -17 points, material constraints -44.8 points, integrity pass→warn
  • Claude Sonnet 4.6: main leaderboard -22.5 points, code execution -15.4 points, material constraints -31.1 points
  • GPT-o3: main leaderboard -22.2 points, code execution -19.5 points, material constraints -25.4 points, integrity pass→warn
  • GPT-5.5: main leaderboard -18.2 points, code execution -33.7 points

Signals to Watch

  • Gemini 2.5 Pro: Today's integrity rating is fail (based on today's Smoke data).

When reading Smoke briefings like this, the focus should be on two questions: first, whether a model has exposed the same type of weakness across multiple consecutive days; second, whether the integrity rating has shifted from pass to warn or fail. Large day-to-day swings in execution or constraint scores may come from question sampling or may be early signals of genuine degradation, requiring follow-up runs for confirmation.


Data source: YZ Index | Run #288 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!