Claude Opus 4.7 Ranks First with 82.55 Points: 2026-08-15 Smoke Quick Test Data Briefing

On 2026-08-15, the YZ Index Smoke quick test covered 9 models, with Claude Opus 4.7 ranking first for the day at 82.55 points. Smoke is a daily 10-question quick test suitable for observing short-term signals, and is not equivalent to the Full weekly ranking conclusions.

This Smoke evaluation only covers two main ranking dimensions: code execution and material constraint. The main ranking formula is 0.55 × Code Execution + 0.45 × Material Constraint. Due to the small daily sample size, single-day scores are better suited as monitoring signals rather than long-term conclusions about model capabilities.

Daily Ranking

RankModelMain RankingCode ExecutionMaterial ConstraintIntegrity
#1Claude Opus 4.782.5599.661.7pass
#2Claude Sonnet 4.680.5210056.7pass
#3GPT-o369.027561.7pass
#4Qwen3 Max69.027561.7pass
#5Doubao Pro68.527560.6pass
#6GPT-5.564.027550.6pass
#7Grok 4595070pass
#8Gemini 3.1 Pro41.522561.7pass
#9Gemini 2.5 Pro31.262538.9pass

Data Interpretation

In today's YZ Index Smoke quick test, Claude Opus 4.7 ranked first with a main ranking score of 82.55. Its combination of 99.6 in code execution and 61.7 in material constraint shows a significant edge in the code execution dimension over material constraint. Claude Sonnet 4.6 followed closely with a main ranking of 80.52, with a structure of 100 in code execution and 56.7 in material constraint, likewise showing high code execution and relatively lower material constraint. GPT-o3 and Qwen3 Max tied at 69.02 on the main ranking, both with 75 in code execution and 61.7 in material constraint. Top models generally rely on code execution to lift their overall scores, while material constraint scores show clear differentiation.

In terms of notable changes, GPT-o3 rose +28.9 on the main ranking, +30 on code execution, and +27.6 on material constraint; Doubao Pro rose +16.2 on the main ranking, +13.3 on code execution, and +19.8 on material constraint; Qwen3 Max rose +15.7 on the main ranking and +30 on code execution; Claude Sonnet 4.6 rose +12.4 on the main ranking and +30 on code execution, but fell -9.1 on material constraint; Claude Opus 4.7 rose +10.3 on the main ranking and +29.6 on code execution, but fell -13.2 on material constraint. Most of these models saw substantial gains in the code execution dimension. GPT-5.5's sharp drop of -15.2 points in material constraint is an anomalous signal that may stem from question sampling fluctuations or could be a sign of genuine degradation, requiring confirmation in subsequent runs. DeepSeek V4 Pro and GLM-4.6 had incomplete data due to the missing execution dimension and API failures/timeouts, so they are not included in this period's ranking and also require re-run verification.

The Smoke quick test is a small-sample single-day signal. The above score structure and fluctuations only reflect that day's results, with measured wording and no long-term inferences made.

Key Changes

  • GPT-o3: Main ranking up 28.9 points, code execution +30, material constraint +27.6
  • Doubao Pro: Main ranking up 16.2 points, code execution +13.3, material constraint +19.8
  • Qwen3 Max: Main ranking up 15.7 points, code execution +30
  • Claude Sonnet 4.6: Main ranking up 12.4 points, code execution +30, material constraint -9.1
  • Claude Opus 4.7: Main ranking up 10.3 points, code execution +29.6, material constraint -13.2

Signals to Watch

  • GPT-5.5: Material constraint plunged -15.2 points
  • DeepSeek V4 Pro: Incomplete data (missing execution dimension, API failure/timeout), auto re-run initiated, not included in this period's ranking
  • GLM-4.6: Incomplete data (missing execution dimension, API failure/timeout), auto re-run initiated, not included in this period's ranking

When reading Smoke briefings like this, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large single-day shifts in execution or constraint scores may stem from question sampling or could be early signs of genuine degradation, requiring confirmation in subsequent runs.


Data source: YZ Index | Run #279 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!