Claude Opus 4.7, GPT-5.5, GPT-6 Astra, and GPT-6.1 Sol Tie at 86.25: 2026-10-02 Smoke Quick Test Data Brief

The 2026-10-02 YZ Index Smoke quick test covered 15 models, with Claude Opus 4.7, GPT-5.5, GPT-6 Astra, and GPT-6.1 Sol tying for first place on the day at 86.25. Smoke is a daily 10-question quick test, suitable for observing short-term signals, and is not equivalent to conclusions from the Full weekly leaderboard.

This Smoke evaluation covered only two main leaderboard dimensions: code execution and material constraints. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraints. Because the daily sample size is small, single-day scores are better treated as monitoring signals rather than long-term conclusions about model capability.

Daily Ranking

RankModelMainCode ExecutionMaterial ConstraintsIntegrity
#1Claude Opus 4.786.2575100pass
#2GPT-5.586.2575100pass
#3GPT-6 Astra86.2575100pass
#4GPT-6.1 Sol86.2575100pass
#5Gemini 2.5 Pro84.5571.9100pass
#6GPT-6 Sol847595pass
#7Grok 483.6770.3100pass
#8Claude Sonnet 4.681.757590pass
#9Doubao Pro81.68775pass
#10GPT-6 Luna79.57585warn
#11GPT-o372.4270.375pass
#12Qwen3 Max71.0571.970pass
#13DeepSeek V4 Pro67.9441.7100pass
#14GLM-4.667.9441.7100fail
#15Gemini 3.1 Pro57.6243.475pass

Data Interpretation

Today’s top four models on the main leaderboard—Claude Opus 4.7, GPT-5.5, GPT-6 Astra, and GPT-6.1 Sol—all show the same structure: 75 in code execution and 100 in material constraints, with a shared main score of 86.25. This combination reflects the supporting effect of a perfect material constraints score on the overall ranking, while code execution remains in the 75 range. Gemini 2.5 Pro scored 84.55 on the main leaderboard with 71.9 in code execution and 100 in material constraints, showing that a perfect material constraints score can still sustain a relatively high position. Doubao Pro scored 81.6 on the main leaderboard, with 87 in code execution and 75 in material constraints, indicating a trade-off between its code execution advantage and its weaker material constraints performance.

Compared with the previous run under the same methodology, Gemini 3.1 Pro fell by 23.6 points on the main leaderboard, 26.6 points in code execution, and 20 points in material constraints. GLM-4.6 fell by 16.1 points on the main leaderboard and 33.3 points in code execution, while its integrity rating changed from pass to fail. GPT-5.5 rose by 16 points on the main leaderboard and 25 points in code execution. Doubao Pro fell by 13.9 points on the main leaderboard and 22 points in material constraints, while DeepSeek V4 Pro fell by 13.3 points on the main leaderboard and 28.3 points in code execution. These changes may stem from sampling variation in the questions, or they may reflect genuine single-day performance differences; follow-up runs are needed for confirmation.

The Smoke test is a small-sample, single-day signal. The current data only show that leading models achieved a main score of 86.25 under a specific combination of code execution and material constraints, while some models showed notable score fluctuations. Interpretation should remain restrained and avoid inferring long-term trends.

Key Changes

  • Gemini 3.1 Pro: main leaderboard down 23.6 points, code execution -26.6 points, material constraints -20 points
  • GLM-4.6: main leaderboard down 16.1 points, code execution -33.3 points, material constraints +5 points, integrity pass→fail
  • GPT-5.5: main leaderboard up 16 points, code execution +25 points, material constraints +5 points
  • Doubao Pro: main leaderboard down 13.9 points, code execution -7.3 points, material constraints -22 points
  • DeepSeek V4 Pro: main leaderboard down 13.3 points, code execution -28.3 points, material constraints +5 points

Signals to Watch

  • GLM-4.6: today’s integrity rating is fail (based on the day’s Smoke data).

When reading this type of Smoke brief, the focus should be on two questions: first, whether a given model exposes the same type of weakness across multiple consecutive days; second, whether its integrity rating moves from pass into warn or fail. Large single-day changes in execution or constraint scores may come from question sampling, or they may be early signals of real degradation, and require follow-up runs for confirmation.


Data source: YZ Index | Run #356 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!