Claude Opus 4.7 Tops with 93.08 Points: 2026-08-30 Smoke Quick Test Data Briefing

On 2026-08-30, the YZ Index Smoke quick test covered 11 models, with Claude Opus 4.7 ranking first for the day with 93.08 points. Smoke is a daily 10-question quick test suited for observing short-term signals, and is not equivalent to Full weekly ranking conclusions.

This Smoke evaluation only covers two main leaderboard dimensions: code execution and material constraint. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraint. Due to the small daily sample size, single-day scores are more suitable as monitoring signals rather than long-term conclusions about model capabilities.

Daily Ranking

RankModelLeaderboardCode ExecutionMaterial ConstraintIntegrity
#1Claude Opus 4.793.0891.595pass
#2GLM-4.688.5891.585pass
#3Grok 488.1490.785pass
#4DeepSeek V4 Pro88.0891.583.9pass
#5Doubao Pro85.8990.780pass
#6Gemini 3.1 Pro85.8391.578.9pass
#7Gemini 2.5 Pro85.7394.575pass
#8Qwen3 Max80.0791.566.1pass
#9Claude Sonnet 4.670.3366.575pass
#10GPT-o369.5666.573.3pass
#11GPT-5.563.0866.558.9pass

Data Interpretation

Among today's top five models on the main leaderboard, Claude Opus 4.7 achieved a leaderboard score of 93.08 with a code execution score of 91.5 and a material constraint score of 95, the highest material constraint score among all models. GLM-4.6 and DeepSeek V4 Pro also scored 91.5 in code execution, but their material constraint scores were 85 and 83.9 respectively, resulting in leaderboard scores of 88.58 and 88.08. Gemini 2.5 Pro scored 94.5 in code execution but only 75 in material constraint, landing at 85.73 on the leaderboard, showing that the combination of strengths in code execution versus material constraint directly affects overall ranking.

GPT-o3 dropped 13.7 points on the leaderboard, with code execution down 8.5 points and material constraint down 20 points. Claude Sonnet 4.6 dropped 12.9 points on the leaderboard, with code execution down 8.5 points and material constraint down 18.3 points. Qwen3 Max dropped 27.2 points in material constraint. Gemini 3.1 Pro dropped 10.8 points on the leaderboard, with code execution down 7.8 points and material constraint down 14.4 points. These changes may stem from question sampling fluctuations or could reflect genuine single-day performance differences, requiring follow-up runs under the same methodology for confirmation.

As for anomalous signals, Grok 4 plunged 8.8 points on the leaderboard, Gemini 3.1 Pro plunged 10.8 points, Qwen3 Max plunged 11.8 points, Claude Sonnet 4.6 plunged 12.9 points, GPT-o3 plunged 13.7 points, and GPT-5.5 plunged 16.7 points in material constraint. The Smoke test provides small-sample single-day signals; these anomalies require more run data to verify their stability.

Key Changes

  • GPT-o3: Leaderboard down 13.7 points, code execution -8.5 points, material constraint -20 points
  • Claude Sonnet 4.6: Leaderboard down 12.9 points, code execution -8.5 points, material constraint -18.3 points
  • Qwen3 Max: Leaderboard down 11.8 points, material constraint -27.2 points
  • Gemini 3.1 Pro: Leaderboard down 10.8 points, code execution -7.8 points, material constraint -14.4 points
  • Claude Opus 4.7: Leaderboard up 9.8 points, code execution +16.5 points

Signals to Watch

  • Grok 4: Leaderboard plunged -8.8 points
  • Gemini 3.1 Pro: Leaderboard plunged -10.8 points
  • Qwen3 Max: Leaderboard plunged -11.8 points
  • Claude Sonnet 4.6: Leaderboard plunged -12.9 points
  • GPT-o3: Leaderboard plunged -13.7 points
  • GPT-5.5: Material constraint plunged -16.7 points

When reading this type of Smoke briefing, the focus should be on two questions: first, whether a model has exposed the same type of weakness on multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large single-day swings in execution or constraint scores may come from question sampling or could be early signals of genuine regression, requiring follow-up runs for verification.


Data source: YZ Index | Run #300 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!