GPT-5.5 Tops the List with 86.25 Points: 2026-09-25 Smoke Quick-Test Data Brief

On 2026-09-25, the YZ Index Smoke quick test covered 10 models, with GPT-5.5 taking the top spot for the day at 86.25 points. Smoke is a daily 10-question quick test suited to observing short-term signals; it is not equivalent to the conclusions of the Full weekly leaderboard.

This Smoke evaluation covered only two main leaderboard dimensions: code execution and material constraints. The main leaderboard formula is 0.55 × code execution + 0.45 × material constraints. Because the daily sample size is small, single-day scores are better treated as monitoring signals than as long-term verdicts on model capability.

Daily Rankings

RankModelMain LeaderboardCode ExecutionMaterial ConstraintsIntegrity
#1GPT-5.586.2575100pass
#2Claude Opus 4.786.0374.6100pass
#3Gemini 3.1 Pro81.6966.7100pass
#4Doubao Pro76.9666.789.5pass
#5GPT-o372.550100pass
#6DeepSeek V4 Pro67.9441.7100pass
#7Claude Sonnet 4.667.785089.5pass
#8Gemini 2.5 Pro67.785089.5pass
#9Grok 461.255075pass
#10Qwen3 Max60.2148.175pass

Data Interpretation

In today's YZ Index Smoke quick test, the leading models showed clear divergence in how code execution paired with material constraints. GPT-5.5 led the main leaderboard at 86.25, with code execution at 75 and material constraints at 100; Claude Opus 4.7 followed closely with 86.03 on the main leaderboard, code execution at 74.6 and material constraints at 100; Gemini 3.1 Pro scored 81.69 on the main leaderboard, with code execution at 66.7 and material constraints at 100. All three reached a perfect score on the material constraints dimension, while their code execution scores separated them, indicating how strongly high material constraints support main leaderboard rankings. Doubao Pro scored 76.96 on the main leaderboard, with code execution at 66.7 and material constraints at 89.5, likewise relying on material constraints to hold a mid-table position.

Compared with the previous run under the same methodology, Grok 4's main leaderboard score fell 21.1 points, with code execution down 45.8 points and material constraints up 9.1 points; DeepSeek V4 Pro's main leaderboard score fell 11.4 points, with code execution down 58.3 points and material constraints up 45.9 points; Gemini 2.5 Pro's main leaderboard score fell 10.4 points, with code execution down 45.8 points and material constraints up 32.8 points. These shifts may stem from sampling variation in the questions or may reflect genuine single-day performance differences, and need confirmation from subsequent runs. Claude Sonnet 4.6's main leaderboard score rose 15.9 points, with material constraints up 35.4 points; GPT-5.5's main leaderboard score rose 15.4 points, with material constraints up 34.2 points — equally in need of multiple rounds of data to verify stability.

The Smoke quick test is a small-sample single-day signal. No anomaly signals are currently on record, and the interpretation is limited to an analysis of today's data structure, with no long-term inferences drawn.

Key Changes

  • Grok 4: main leaderboard down 21.1 points, code execution −45.8 points, material constraints +9.1 points
  • Claude Sonnet 4.6: main leaderboard up 15.9 points, material constraints +35.4 points
  • GPT-5.5: main leaderboard up 15.4 points, material constraints +34.2 points
  • DeepSeek V4 Pro: main leaderboard down 11.4 points, code execution −58.3 points, material constraints +45.9 points
  • Gemini 2.5 Pro: main leaderboard down 10.4 points, code execution −45.8 points, material constraints +32.8 points

Signals to Watch

  • No publishable anomaly signals were retained this time.

When reading this kind of Smoke brief, the focus should be on two questions: first, whether a given model exposes the same type of weakness across several consecutive days; second, whether its integrity rating moves from pass into warn or fail. Large single-day swings in execution or constraint scores may come from question sampling, or they may be early signals of genuine degradation, and require confirmation from subsequent runs.


Data source: YZ Index | Run #338 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!