Claude Opus 4.7 Leads with 97.48 Points: 2026-10-07 Smoke Quick-Test Data Brief

On 2026-10-07, the YZ Index Smoke quick test covered 14 models, with Claude Opus 4.7 ranking first that day at 97.48. Smoke is a daily 10-question quick test, suitable for observing short-term signals; it is not equivalent to the conclusions of the weekly Full leaderboard.

This Smoke evaluation covered only the two main leaderboard dimensions of code execution and material constraints. The main leaderboard formula is 0.55 × code execution + 0.45 × material constraints. Because the daily sample size is small, single-day scores are better used as monitoring signals rather than long-term conclusions about model capabilities.

Daily Ranking

RankModelMain leaderboardCode executionMaterial constraintsIntegrity
#1Claude Opus 4.797.4810094.4pass
#2Qwen3 Max95.1297.592.2pass
#3GPT-6.1 Sol91.3210080.7pass
#4GPT-o388.810075.1pass
#5GPT-5.588.4410074.3pass
#6Grok 483.9896.868.3pass
#7GPT-6 Luna83.697594.3pass
#8GPT-6 Sol79.7510055pass
#9GPT-6 Astra77.577580.7pass
#10Gemini 3.1 Pro76.9878.675pass
#11Claude Sonnet 4.669.477562.7pass
#12Doubao Pro64.0474.351.5pass
#13DeepSeek V4 Pro61.255075pass
#14Gemini 2.5 Pro56.9252.162.8pass

Data Interpretation

Today's top three models on the main leaderboard show different emphases in the pairing of code execution and material constraints. Claude Opus 4.7 scored 100 in code execution and 94.4 in material constraints for a main leaderboard score of 97.48; Qwen3 Max scored 97.5 in code execution and 92.2 in material constraints for 95.12; GPT-6.1 Sol scored 100 in code execution and 80.7 in material constraints for 91.32. GPT-6 Luna scored 75 in code execution and 94.3 in material constraints, with a main leaderboard score of 83.69, showing a structural characteristic of relatively stronger material constraints. Among the models with larger changes, GPT-6.1 Sol rose 35.3 points on the main leaderboard, 50 points in code execution, and 17.4 points in material constraints; Claude Opus 4.7 rose 30.2 points on the main leaderboard, 50 points in code execution, and 6.1 points in material constraints; Qwen3 Max rose 30.1 points on the main leaderboard and 51.5 points in code execution.

Regarding anomalous signals, GPT-6 Sol's material constraints plunged by -33.3 points, Claude Sonnet 4.6's material constraints plunged by -30.6 points, and Doubao Pro's material constraints plunged by -16.8 points. These single-day numerical fluctuations may stem from differences in question sampling, or they may indicate temporary changes in models on the material constraints dimension; follow-up runs using the same methodology are needed to confirm. GLM-4.6 did not participate in the ranking this period because of incomplete data due to an API failure. The Smoke quick test is a small-sample, single-day signal; the above observations reflect only that day's results.

Main Changes

  • GPT-6.1 Sol: main leaderboard up 35.3 points, code execution +50 points, material constraints +17.4 points
  • Claude Opus 4.7: main leaderboard up 30.2 points, code execution +50 points, material constraints +6.1 points
  • Qwen3 Max: main leaderboard up 30.1 points, code execution +51.5 points
  • Gemini 3.1 Pro: main leaderboard up 25.6 points, code execution +36.9 points, material constraints +11.7 points
  • GPT-o3: main leaderboard up 10.8 points, code execution +25 points, material constraints -6.6 points

Signals to Watch

  • GPT-6 Sol: material constraints plunged -33.3 points
  • Claude Sonnet 4.6: material constraints plunged -30.6 points
  • Doubao Pro: material constraints plunged -16.8 points
  • GLM-4.6: incomplete data (missing one evaluation dimension, API failure/timeout); automatic rerun has been initiated; does not participate in this period's ranking

When reading this kind of Smoke brief, focus on two questions: first, whether a model exposes the same type of weakness on consecutive days; second, whether its integrity rating moves from pass to warn or fail. Large single-day changes in execution or constraint scores may come from question sampling or may be early signals of real degradation, requiring follow-up runs to verify.


Data source: YZ Index | Run #364 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!