Claude Opus 4.7, GPT-6 Astra, and GPT-6.1 Sol Tie at 79.79: 2026-10-11 Smoke Quick-Test Data Brief

The 2026-10-11 YZ Index Smoke quick test covered 13 models, with Claude Opus 4.7, GPT-6 Astra, and GPT-6.1 Sol tied for first place that day at 79.79 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals; it is not equivalent to the conclusions of the weekly Full leaderboard.

This Smoke evaluation covered only two main leaderboard dimensions: code execution and material constraints. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraints. Because the daily sample size is small, single-day scores are better treated as monitoring signals rather than as long-term verdicts on model capability.

Daily Rankings

RankModelMain LeaderboardCode ExecutionMaterial ConstraintsIntegrity
#1Claude Opus 4.779.797289.3pass
#2GPT-6 Astra79.797289.3pass
#3GPT-6.1 Sol79.797289.3pass
#4Claude Sonnet 4.674.527277.6pass
#5Gemini 3.1 Pro68.547264.3pass
#6Gemini 2.5 Pro66.044789.3pass
#7GPT-6 Sol66.044789.3pass
#8Grok 464.927552.6pass
#9GPT-6 Luna62.17250pass
#10Doubao Pro54.794764.3pass
#11GPT-o354.794764.3pass
#12DeepSeek V4 Pro49.524752.6pass
#13GPT-5.548.354750pass

Data Interpretation

Today's top three on the main leaderboard — Claude Opus 4.7, GPT-6 Astra, and GPT-6.1 Sol — all scored 79.79, with code execution at 72 and material constraints at 89.3. The three are structurally identical, showing that while they maintain a high level on the material constraints dimension, their code execution scores fall in the same range. Fourth-place Claude Sonnet 4.6 followed at 74.52, with code execution also at 72 but material constraints dropping to 77.6, reflecting the impact of relatively weaker material constraints on the overall ranking. Gemini 3.1 Pro recorded 72 for code execution and 64.3 for material constraints, for a main leaderboard score of 68.54, while Grok 4 scored 64.92 with 75 for code execution and 52.6 for material constraints — illustrating the different ways strengths in code execution and material constraints combined across models in the day's score distribution.

DeepSeek V4 Pro's main leaderboard score fell by 47.4 points, with code execution down 53 points and material constraints down 40.6 points; GPT-5.5 fell 38.6 points on the main leaderboard, with code execution down 53 points and material constraints down 20.9 points; GPT-o3 fell 35 points on the main leaderboard, with code execution down 51.5 points and material constraints down 14.8 points; Doubao Pro fell 32.1 points on the main leaderboard, with code execution down 53 points and material constraints down 6.6 points; GPT-6 Luna fell 26.7 points on the main leaderboard, with code execution down 28 points and material constraints down 25 points. These significant changes may stem from sampling fluctuations in the day's questions, or they may reflect genuine performance degradation, and require follow-up runs using the same methodology to confirm. The Smoke quick test is a small-sample, single-day signal; the observations above are for reference that day only and do not constitute a basis for long-term judgment.

Key Changes

  • DeepSeek V4 Pro: main leaderboard down 47.4 points, code execution −53 points, material constraints −40.6 points
  • GPT-5.5: main leaderboard down 38.6 points, code execution −53 points, material constraints −20.9 points
  • GPT-o3: main leaderboard down 35 points, code execution −51.5 points, material constraints −14.8 points
  • Doubao Pro: main leaderboard down 32.1 points, code execution −53 points, material constraints −6.6 points
  • GPT-6 Luna: main leaderboard down 26.7 points, code execution −28 points, material constraints −25 points

Signals to Watch

  • No publishable anomalous signals were retained in this run.

When reading this kind of Smoke brief, the focus should be on two questions: first, whether a given model shows the same type of weakness on consecutive days; second, whether its integrity rating moves from pass into warn or fail. Large single-day changes in execution or constraint scores may come from question sampling, or they may be early signals of genuine degradation, and require follow-up runs to confirm.


Data source: YZ Index | Run #371 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!