Claude Opus 4.7 Tops with 98.65: 2026-10-03 Smoke Quick Test Data Brief

On 2026-10-03, the YZ Index Smoke quick test covered 15 models, with Claude Opus 4.7 ranking first for the day at 98.65. Smoke is a daily 10-question quick test, suited to observing short-term signals; it is not equivalent to Full weekly ranking conclusions.

This Smoke evaluation covered only two main leaderboard dimensions: Code Execution and Material Constraints. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraints. Because the daily sample size is small, single-day scores are better used as monitoring signals rather than long-term conclusions about model capabilities.

Daily Ranking

RankModelMain LeaderboardCode ExecutionMaterial ConstraintsIntegrity
#1Claude Opus 4.798.6510097pass
#2GPT-o392.7110083.8pass
#3Gemini 2.5 Pro87.410072pass
#4Gemini 3.1 Pro87.410072pass
#5GPT-6 Sol85.267597.8pass
#6Grok 484.97597pass
#7DeepSeek V4 Pro83.917594.8pass
#8Claude Sonnet 4.679.6573.886.8pass
#9GPT-6.1 Sol77.0758.3100pass
#10GPT-6 Astra74.7358.394.8pass
#11Doubao Pro72.2883.358.8pass
#12GPT-5.568.7458.381.5pass
#13GPT-6 Luna68.7458.381.5pass
#14Qwen3 Max66.457556pass
#15GLM-4.638.755025warn

Data Interpretation

In today's YZ Index Smoke quick test, Claude Opus 4.7 ranked first with a main leaderboard score of 98.65; its combination of 100 in Code Execution and 97 in Material Constraints shows balanced but strong performance in both. GPT-o3 scored 92.71 on the main leaderboard, also with 100 in Code Execution but 83.8 in Material Constraints, reflecting a structure led by Code Execution. Gemini 2.5 Pro and Gemini 3.1 Pro both scored 87.4 on the main leaderboard, with a combination of 100 in Code Execution and 72 in Material Constraints, highlighting an advantage in Code Execution. GPT-6 Sol scored 85.26 on the main leaderboard, with 75 in Code Execution and 97.8 in Material Constraints, showing a complementary characteristic of stronger Material Constraints, while Grok 4 and DeepSeek V4 Pro maintained a similar pairing of 75 in Code Execution and 97 and 94.8 in Material Constraints, respectively.

Gemini 3.1 Pro's main leaderboard score rose 29.8 compared with the previous comparable run, and Code Execution rose 56.6, indicating a clear improvement on the Code Execution side; GPT-o3's main leaderboard score rose 20.3, with Code Execution up 29.7 and Material Constraints up 8.8; DeepSeek V4 Pro's main leaderboard score rose 16, with Code Execution up 33.3. GLM-4.6's main leaderboard score fell 29.2, Material Constraints fell 75, and Integrity changed from fail to warn; GPT-5.5's main leaderboard score fell 17.5, with Code Execution down 16.7 and Material Constraints down 18.5.

In terms of anomalous signals, Gemini 2.5 Pro's Material Constraints plunged 28, GPT-6.1 Sol's main leaderboard score plunged 9.2, GPT-6 Astra's main leaderboard score plunged 11.5, Doubao Pro's main leaderboard score plunged 9.3, GPT-5.5's main leaderboard score plunged 17.5, GPT-6 Luna's main leaderboard score plunged 10.8, and GLM-4.6's main leaderboard score plunged 29.2; these may stem from question sampling fluctuations or real degradation and need follow-up runs to verify. As a small-sample single-day signal, Smoke's observations above only reflect characteristics of that day's data structure.

Major Changes

  • Gemini 3.1 Pro: main leaderboard up 29.8 points, Code Execution +56.6 points
  • GLM-4.6: main leaderboard down 29.2 points, Code Execution +8.3 points, Material Constraints -75 points, Integrity fail→warn
  • GPT-o3: main leaderboard up 20.3 points, Code Execution +29.7 points, Material Constraints +8.8 points
  • GPT-5.5: main leaderboard down 17.5 points, Code Execution -16.7 points, Material Constraints -18.5 points
  • DeepSeek V4 Pro: main leaderboard up 16 points, Code Execution +33.3 points, Material Constraints -5.2 points

Signals to Watch

  • Gemini 2.5 Pro: Material Constraints plunged -28 points
  • GPT-6.1 Sol: main leaderboard plunged -9.2 points
  • GPT-6 Astra: main leaderboard plunged -11.5 points
  • Doubao Pro: main leaderboard plunged -9.3 points
  • GPT-5.5: main leaderboard plunged -17.5 points
  • GPT-6 Luna: main leaderboard plunged -10.8 points
  • GLM-4.6: main leaderboard plunged -29.2 points

When reading this kind of Smoke brief, the focus should be on two questions: first, whether a model exposes the same type of weakness over multiple consecutive days; second, whether its Integrity rating moves from pass into warn or fail. Large single-day changes in execution or constraint scores may come from question sampling, or they may be early signals of real degradation, requiring follow-up runs for verification.


Data source: YZ Index | Run #358 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!