Claude Sonnet 4.6 Tops with 96.98 Points: 2026-09-23 Smoke Quick-Test Data Brief

On 2026-09-23, the YZ Index Smoke quick test covered 10 models, with Claude Sonnet 4.6 ranking first for the day at 96.98 points. Smoke is a daily 10-question quick test, suited to observing short-term signals; it is not equivalent to the conclusions of the Full weekly leaderboard.

This Smoke evaluation covered only two main leaderboard dimensions: code execution and material constraints. The main leaderboard formula is 0.55 × code execution + 0.45 × material constraints. Because the daily sample size is small, single-day scores are better used as monitoring signals than as long-term conclusions about model capability.

Daily Ranking

RankingModelMain LeaderboardCode ExecutionMaterial ConstraintsIntegrity
#1Claude Sonnet 4.696.9894.5100pass
#2Claude Opus 4.796.7694.1100pass
#3Gemini 2.5 Pro95.1993.797pass
#4GPT-5.591.8989.594.8pass
#5Qwen3 Max91.2685.997.8pass
#6Grok 480.8969.594.8pass
#7GPT-o380.4864.5100pass
#8Gemini 3.1 Pro79.1364.597pass
#9DeepSeek V4 Pro78.1464.594.8pass
#10Doubao Pro73.564.584.5pass

Data Interpretation

In today's YZ Index Smoke quick test, leading models showed a balanced high-end pattern in combining code execution and material constraints. Claude Sonnet 4.6 ranked first on the main leaderboard at 96.98, with code execution 94.5 and material constraints 100; Claude Opus 4.7 had main leaderboard 96.76, code execution 94.1, material constraints 100; Gemini 2.5 Pro had main leaderboard 95.19, code execution 93.7, material constraints 97. All three maintained Integrity pass, and the material constraints dimension generally approached or reached 100, while code execution was in the 93.7 to 94.5 range, forming a structurally complementary advantage of strengths.

Several models showed notable changes under the same measurement basis. Gemini 2.5 Pro rose 22.7 points on the main leaderboard, with code execution +21.7 points and material constraints +23.9 points; Qwen3 Max rose 12.8 points on the main leaderboard, with code execution +13.9 points and material constraints +11.5 points; DeepSeek V4 Pro rose 12.1 points on the main leaderboard, with code execution -7.5 points and material constraints +36.1 points. Doubao Pro fell 17.3 points on the main leaderboard, with code execution -31.8 points; GPT-o3 fell 15.5 points on the main leaderboard, with code execution -32.5 points and material constraints +5.2 points. These changes stem from small-sample sampling of 10 questions on a single day and require follow-up runs to distinguish fluctuation from genuine degradation.

Anomalous signals were concentrated in several models. Grok 4 plunged -10 points on the main leaderboard, GPT-o3 plunged -15.5 points on the main leaderboard, Gemini 3.1 Pro plunged -31.8 points in code execution, and Doubao Pro plunged -17.3 points on the main leaderboard. GLM-4.6 had incomplete data due to an API failure and did not participate in the ranking. These are all single-day Smoke signals; the magnitude of change may be affected by question sampling, and multiple rounds of retesting are needed to confirm stability.

Major Changes

  • Gemini 2.5 Pro: main leaderboard up 22.7 points, code execution +21.7 points, material constraints +23.9 points
  • Doubao Pro: main leaderboard down 17.3 points, code execution -31.8 points
  • GPT-o3: main leaderboard down 15.5 points, code execution -32.5 points, material constraints +5.2 points
  • Qwen3 Max: main leaderboard up 12.8 points, code execution +13.9 points, material constraints +11.5 points
  • DeepSeek V4 Pro: main leaderboard up 12.1 points, code execution -7.5 points, material constraints +36.1 points

Signals to Watch

  • Grok 4: main leaderboard plunged -10 points
  • GPT-o3: main leaderboard plunged -15.5 points
  • Gemini 3.1 Pro: code execution plunged -31.8 points
  • Doubao Pro: main leaderboard plunged -17.3 points
  • GLM-4.6: incomplete data (missing several evaluation dimensions due to API failure/timeout), has entered automatic rerun, and is not ranked in this period

When reading this kind of Smoke brief, the focus should be on two questions: first, whether a model has repeatedly shown the same type of weakness over multiple consecutive days; second, whether the integrity rating moves from pass to warn or fail. Large single-day changes in execution or constraint scores may come from question sampling, or they may be early signals of genuine degradation, requiring follow-up runs to confirm.


Data source: YZ Index | Run #335 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!