GPT-5.5 Tops with 93.16 Points: 2026-09-28 Smoke Quick Test Data Brief

2026-09-28 YZ Index Smoke quick test covered 11 models, with GPT-5.5 ranking first for the day at 93.16 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals; it is not equivalent to the conclusions of the weekly Full leaderboard.

This Smoke evaluation covers only two Main Leaderboard dimensions: Code Execution and Material Constraints. The Main Leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraints. Because the daily sample size is small, single-day scores are better treated as monitoring signals rather than long-term conclusions about model capability.

Daily Ranking

RankingModelMain LeaderboardCode ExecutionMaterial ConstraintsIntegrity
#1GPT-5.593.169590.9pass
#2Claude Sonnet 4.691.8110081.8pass
#3Grok 491.8110081.8pass
#4Qwen3 Max91.7892.590.9warn
#5Claude Opus 4.782.167590.9pass
#6GPT-o380.7872.590.9pass
#7DeepSeek V4 Pro70.917565.9warn
#8Gemini 3.1 Pro68.6378.356.8pass
#9Doubao Pro68.167065.9warn
#10Gemini 2.5 Pro64.067056.8pass
#11GLM-4.651.912090.9pass

Data Interpretation

Among today's top three on the Main Leaderboard, GPT-5.5 scored 93.16 with a combination of Code Execution 95 and Material Constraints 90.9, while Claude Sonnet 4.6 and Grok 4 both tied at 91.81 with Code Execution 100 and Material Constraints 81.8, showing that leading models differ in their relative strengths across the two capabilities. Qwen3 Max followed closely with Code Execution 92.5, Material Constraints 90.9, and a Main Leaderboard score of 91.78, while Claude Opus 4.7 and GPT-o3 both showed a pattern of lower Code Execution with Material Constraints 90.9, with Main Leaderboard scores of 82.16 and 80.78, respectively.

On anomalous signals, Claude Opus 4.7's Main Leaderboard score plunged by 13.9 points, GPT-o3's Main Leaderboard score plunged by 15.2 points, Doubao Pro's Main Leaderboard score plunged by 14.8 points and its Integrity rating shifted from pass to warn, and GLM-4.6's Code Execution plunged by 47 points. These single-day changes may stem from question sampling fluctuations or may reflect genuine degradation, and need to be confirmed by follow-up runs using the same methodology.

The Smoke quick test has a limited sample size; the above figures are only observations for the day and should not be extrapolated to a model's overall performance.

Key Changes

  • GPT-o3: Main Leaderboard down 15.2 points, Code Execution -24.5 points
  • Doubao Pro: Main Leaderboard down 14.8 points, Code Execution -22 points, Material Constraints -6.1 points, Integrity pass→warn
  • GPT-5.5: Main Leaderboard up 14.4 points, Code Execution +23 points
  • Claude Opus 4.7: Main Leaderboard down 13.9 points, Code Execution -22 points
  • Grok 4: Main Leaderboard down 5.8 points, Material Constraints -13 points

Signals to Watch

  • Claude Opus 4.7: Main Leaderboard plunges by 13.9 points
  • GPT-o3: Main Leaderboard plunges by 15.2 points
  • Doubao Pro: Main Leaderboard plunges by 14.8 points
  • GLM-4.6: Code Execution plunges by 47 points

When reading this type of Smoke brief, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether its Integrity rating has moved from pass into warn or fail. Large single-day changes in execution or constraint scores may come from question sampling or may be early signals of genuine degradation, and need to be verified by follow-up runs.


Data source: YZ Index | Run #341 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!