Claude Opus 4.7 Tops with 95.63 Points: 2026-10-05 Smoke Quick-Test Data Brief

The 2026-10-05 YZ Index Smoke quick test covered 14 models, with Claude Opus 4.7 ranking first for the day at 95.63 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals and not equivalent to Full weekly leaderboard conclusions.

This Smoke evaluation covered only two main leaderboard dimensions: code execution and material constraints. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraints. Because the daily sample size is small, single-day scores are better suited as monitoring signals rather than long-term conclusions about model capability.

Daily Rankings

RankModelMain LeaderboardCode ExecutionMaterial ConstraintsIntegrity
#1Claude Opus 4.795.6394.597pass
#2GPT-o390.3694.585.3pass
#3GPT-6 Sol89.3794.583.1pass
#4Doubao Pro88.691.385.3pass
#5Gemini 2.5 Pro88.691.385.3pass
#6DeepSeek V4 Pro83.1191.373.1pass
#7Qwen3 Max78.4265.594.2pass
#8Gemini 3.1 Pro76.6169.585.3pass
#9GPT-5.575.6269.583.1pass
#10GPT-6 Luna75.6269.583.1pass
#11GPT-6 Astra73.4669.578.3pass
#12Grok 472.66483.1pass
#13GPT-6.1 Sol71.1269.573.1pass
#14Claude Sonnet 4.664.2144.588.3pass

Data Interpretation

In today's YZ Index Smoke quick test, Claude Opus 4.7 ranked first on the main leaderboard with 95.63; its pairing of 94.5 in code execution and 97 in material constraints shows a structure with both dimensions relatively high and balanced. GPT-o3 scored 90.36 on the main leaderboard, with 94.5 in code execution and 85.3 in material constraints, also maintaining a high position supported by a high code execution score. Doubao Pro and Gemini 2.5 Pro both scored 88.6 on the main leaderboard, with 91.3 in code execution and 85.3 in material constraints, showing a combination in which code execution is slightly stronger than material constraints. Qwen3 Max's main leaderboard score of 78.42 relied on a relatively high 94.2 in material constraints, while its code execution was only 65.5, forming a material-constraint-dominated structure.

GPT-6 Sol rose 28.2 points on the main leaderboard from the previous run, with code execution up 24.7 points and material constraints up 32.5 points; Gemini 2.5 Pro rose 24.9 points on the main leaderboard, with code execution up 21.5 points and material constraints up 29.1 points; GPT-6 Astra rose 20.6 points on the main leaderboard, with code execution up 24.7 points and material constraints up 15.6 points. The score structures of these unusually moving models show that material constraints improved either in tandem with code execution or with a bias toward material constraints. GPT-5.5 rose 15.5 points on the main leaderboard, with code execution down 5.5 points but material constraints up 41.2 points; Claude Opus 4.7 rose 14.1 points on the main leaderboard and material constraints up 34.3 points, reflecting that fluctuations in the material constraints dimension had a clear impact on overall ranking.

On anomalous signals, Grok 4's code execution plunged 31.3 points and Claude Sonnet 4.6's code execution plunged 27.4 points; this may stem from question sampling fluctuations or temporary state changes during a single day's run, and needs follow-up run confirmation. GLM-4.6 did not participate in the ranking due to incomplete data in the judgment dimension, also suggesting that API-related issues may have interfered with results. The Smoke quick test is a small-sample, single-day signal; the above observations are all based on that day's data and no long-term inference is made.

Key Changes

  • GPT-6 Sol: Main leaderboard up 28.2 points, code execution +24.7 points, material constraints +32.5 points
  • Gemini 2.5 Pro: Main leaderboard up 24.9 points, code execution +21.5 points, material constraints +29.1 points
  • GPT-6 Astra: Main leaderboard up 20.6 points, code execution +24.7 points, material constraints +15.6 points
  • GPT-5.5: Main leaderboard up 15.5 points, code execution -5.5 points, material constraints +41.2 points
  • Claude Opus 4.7: Main leaderboard up 14.1 points, material constraints +34.3 points

Signals to Watch

  • Grok 4: Code execution plunged -31.3 points
  • Claude Sonnet 4.6: Code execution plunged -27.4 points
  • GLM-4.6: Incomplete data (missing judgment dimension, API failure/timeout), has entered an automatic re-run, and is not included in this period's ranking

When reading this kind of Smoke brief, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether its integrity rating moves from pass to warn or fail. Large single-day changes in execution or constraint scores may come from question sampling, or they may be an early signal of genuine degradation, requiring follow-up run confirmation.


Data source: YZ Index | Run #361 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!