Claude Opus 4.7 Tops Leaderboard with 92.49 Points: 2026-09-04 Smoke Quick Test Data Brief

On September 4, 2026, the YZ Index Smoke quick test covered 11 models, and Claude Opus 4.7 ranked first that day with 92.49 points. Smoke is a daily quick test of 10 questions, suitable for observing short-term signals, and is not equivalent to the conclusions of the Full weekly ranking.

This Smoke evaluation only covers two main leaderboard dimensions: code execution and material constraints. The main leaderboard formula is 0.55 × code execution + 0.45 × material constraints. Due to the small daily sample size, single-day scores are more suitable as monitoring signals rather than long-term conclusions about model capabilities.

Daily Ranking

RankModelMain LeaderboardCode ExecutionMaterial ConstraintsIntegrity
#1Claude Opus 4.792.4910083.3pass
#2Grok 489.3810076.4pass
#3GPT-o382.4797.564.1pass
#4Gemini 2.5 Pro75.457576pass
#5Qwen3 Max75.1596.748.8pass
#6GPT-5.571.677567.6pass
#7DeepSeek V4 Pro67.245088.3pass
#8GLM-4.660.945074.3pass
#9Gemini 3.1 Pro60.657543.1pass
#10Claude Sonnet 4.652.165054.8pass
#11Doubao Pro50.095050.2pass

Data Interpretation

The top three models on today's main leaderboard presented different combinations across the code execution and material constraint metrics. Claude Opus 4.7 achieved a main leaderboard score of 92.49 with 100 in code execution and 83.3 in material constraints; Grok 4 also scored 100 in code execution but 76.4 in material constraints, for a main leaderboard score of 89.38; GPT-o3 scored 97.5 in code execution and 64.1 in material constraints, for a main leaderboard score of 82.47. All three received an integrity rating of pass. DeepSeek V4 Pro scored 50 in code execution and 88.3 in material constraints, for a main leaderboard score of 67.24, reflecting a structural profile in which material constraints are relatively stronger. Gemini 2.5 Pro scored 75 in code execution and 76 in material constraints, for a main leaderboard score of 75.45, sitting at a relatively balanced position across the two metrics.

Multiple models showed notable changes on a comparable basis. Claude Opus 4.7: main leaderboard +33.6 points, code execution +25 points, material constraints +44 points; Grok 4: main leaderboard +27.4 points, code execution +25 points, material constraints +30.3 points; Gemini 2.5 Pro: main leaderboard +27.9 points, code execution +25 points, material constraints +31.5 points; DeepSeek V4 Pro: main leaderboard +22.1 points, material constraints +49 points. Doubao Pro: main leaderboard -29 points, code execution -50 points. These single-day swings should be viewed in light of Smoke's small-sample nature; they may stem from question-sampling fluctuation, or may reflect genuine performance changes, and all require confirmation through subsequent runs.

Anomalous signals were concentrated in several models. GPT-5.5's code execution plunged by 25 points, GLM-4.6's integrity rating shifted from fail to Fail, Claude Sonnet 4.6's main leaderboard score plunged by 13.9 points, and Doubao Pro's main leaderboard score plunged by 29 points. Such changes are common in small-sample single-day tests and should not yet be viewed as long-term regression; they serve only as observation signals, and further validation through multiple rounds of repeated testing is recommended.

Key Changes

  • Claude Opus 4.7: main leaderboard up 33.6 points, code execution +25 points, material constraints +44 points
  • Doubao Pro: main leaderboard down 29 points, code execution -50 points
  • Gemini 2.5 Pro: main leaderboard up 27.9 points, code execution +25 points, material constraints +31.5 points
  • Grok 4: main leaderboard up 27.4 points, code execution +25 points, material constraints +30.3 points
  • DeepSeek V4 Pro: main leaderboard up 22.1 points, material constraints +49 points

Signals to Watch

  • GPT-5.5: code execution plunged by 25 points
  • GLM-4.6: integrity rating downgraded to Fail (fail→pass)
  • Claude Sonnet 4.6: main leaderboard plunged by 13.9 points
  • Doubao Pro: main leaderboard plunged by 29 points

When reading Smoke briefs of this kind, the focus should be on two questions: first, whether a given model has exposed the same type of weakness on multiple consecutive days; second, whether the integrity rating has moved from pass into warn or fail. Large single-day swings in execution or constraint scores may come from question sampling, or may be early signals of genuine degradation, and require review in subsequent runs.


Data source: YZ Index | Run #308 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!