Claude Sonnet 4.6 Leads with 87.41: 2026-10-08 Smoke Quick Test Data Brief

2026-10-08 YZ Index Smoke quick test covered 15 models, with Claude Sonnet 4.6 ranking first for the day at 87.41 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals, and is not equivalent to Full weekly leaderboard conclusions.

This Smoke evaluation only covers two main leaderboard dimensions: Code Execution and Material Constraints. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraints. Because the daily sample size is small, single-day scores are more suitable as monitoring signals rather than long-term conclusions about model capabilities.

Daily Ranking

RankModelMain LeaderboardCode ExecutionMaterial ConstraintsIntegrity
#1Claude Sonnet 4.687.4198.773.6pass
#2GPT-6 Luna83.197593.2pass
#3Claude Opus 4.779.17584.1pass
#4GPT-o375.637576.4pass
#5GPT-5.574.377573.6pass
#6Grok 468.3270.365.9pass
#7GPT-6 Sol63.757550pass
#8Qwen3 Max62.367151.8pass
#9GPT-6 Astra61.255075pass
#10GPT-6.1 Sol61.255075pass
#11Doubao Pro60.625073.6pass
#12Gemini 2.5 Pro59.9148.773.6pass
#13Gemini 3.1 Pro58.825069.6pass
#14DeepSeek V4 Pro55.692593.2pass
#15GLM-4.629.12534.1pass

Data Interpretation

From the score structure, Claude Sonnet 4.6 achieved a Main Leaderboard score of 87.41 with a combination of 98.7 in Code Execution and 73.6 in Material Constraints, showing a standout advantage in Code Execution; GPT-6 Luna achieved a Main Leaderboard score of 83.19 with 75 in Code Execution and 93.2 in Material Constraints, with a higher Material Constraints score, forming a different strength pairing. Claude Opus 4.7, with 75 in Code Execution and 84.1 in Material Constraints, scored 79.1 on the Main Leaderboard, similarly reflecting a relatively balanced Material Constraints profile. Among top models, such pairings differ noticeably. GPT-6 Astra and GPT-6.1 Sol both have 50 in Code Execution and 75 in Material Constraints, with identical Main Leaderboard scores of 61.25.

In terms of notable changes, Qwen3 Max dropped 32.8 points on the Main Leaderboard, 26.5 points in Code Execution, and 40.4 points in Material Constraints; GPT-6.1 Sol dropped 30.1 points on the Main Leaderboard and 50 points in Code Execution; GLM-4.6 dropped 29.7 points on the Main Leaderboard and 65.9 points in Material Constraints; Claude Opus 4.7 dropped 18.4 points on the Main Leaderboard and 25 points in Code Execution; Gemini 3.1 Pro dropped 18.2 points on the Main Leaderboard and 28.6 points in Code Execution. These fluctuations may stem from question sampling variability, or they may reflect genuine model performance degradation and require follow-up runs to verify.

The Smoke quick test is a small-sample, single-day signal; the above analysis is based only on today's data, with restrained wording and no long-term judgments.

Key Changes

  • Qwen3 Max: Main Leaderboard down 32.8 points, Code Execution -26.5 points, Material Constraints -40.4 points
  • GPT-6.1 Sol: Main Leaderboard down 30.1 points, Code Execution -50 points, Material Constraints -5.7 points
  • GLM-4.6: Main Leaderboard down 29.7 points, Material Constraints -65.9 points, Integrity warn→pass
  • Claude Opus 4.7: Main Leaderboard down 18.4 points, Code Execution -25 points, Material Constraints -10.3 points
  • Gemini 3.1 Pro: Main Leaderboard down 18.2 points, Code Execution -28.6 points, Material Constraints -5.4 points

Signals to Watch

  • No publishable anomaly signals were retained this time.

When reading this type of Smoke brief, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether its Integrity rating moves from pass to warn or fail. Large single-day changes in execution or constraint scores may come from question sampling, or they may be early signals of genuine degradation and require follow-up runs for verification.


Data source: YZ Index | Run #366 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!