GPT-o3 Tops with 96.01: 2026-09-22 Smoke Quick Test Data Brief

2026-09-22 YZ Index Smoke quick test covered 10 models, with GPT-o3 ranking first for the day at 96.01. Smoke is a daily 10-question quick test, suitable for observing short-term signals; it is not equivalent to conclusions from the full weekly rankings.

This Smoke evaluation covers only two main leaderboard dimensions: code execution and material constraints. The main leaderboard formula is 0.55 × code execution + 0.45 × material constraints. Because the daily sample size is small, single-day scores are better used as monitoring signals rather than long-term conclusions about model capability.

Daily Rankings

RankModelMainCode ExecutionMaterial ConstraintsIntegrity
#1GPT-o396.019794.8pass
#2Claude Sonnet 4.691.249784.2pass
#3Claude Opus 4.791.159784pass
#4Grok 490.8696.384.2pass
#5Doubao Pro90.7796.384pass
#6GPT-5.586.169772.9pass
#7Gemini 3.1 Pro79.2996.358.5pass
#8Qwen3 Max78.447286.3pass
#9Gemini 2.5 Pro72.57273.1pass
#10DeepSeek V4 Pro66.027258.7pass

Data Interpretation

In today's YZ Index Smoke quick test, GPT-o3 took first place with a main score of 96.01 from a combination of 97 in code execution and 94.8 in material constraints; its overall performance under the 0.55/0.45 weighting was outstanding. Claude Sonnet 4.6 and Claude Opus 4.7 also had code execution 97, but their material constraints were 84.2 and 84 respectively, with main scores of 91.24 and 91.15; Grok 4 and Doubao Pro had code execution 96.3 and material constraints 84.2 and 84, ranking immediately behind. Qwen3 Max showed a complementary structure of code execution 72 and material constraints 86.3, for a main score of 78.44.

Among notable movers, DeepSeek V4 Pro had a main score of 66.02, down 29 points from the previous time, with code execution 72 and material constraints 58.7, falling 28 and 30.2 points respectively; Claude Opus 4.7 rose 20.4 points on the main leaderboard, mainly from a 41.4-point increase in code execution; Claude Sonnet 4.6 gained 47 points in code execution and rose 18.7 points on the main leaderboard. Gemini 3.1 Pro fell 15.9 points on the main leaderboard, with material constraints down 30.8 points; GPT-o3 rose 14.6 points on the main leaderboard, with code execution up 22 points.

In terms of anomalous signals, Claude Sonnet 4.6 and Grok 4 both saw material constraints fall 15.8 points, Gemini 2.5 Pro fell 13.1 points on the main leaderboard, and DeepSeek V4 Pro fell 29 points on the main leaderboard; these changes may stem from question sampling fluctuations or genuine degradation and require follow-up run review. As a small-sample single-day signal, the current Smoke data is for reference only.

Key Changes

  • DeepSeek V4 Pro: Main down 29 points, code execution -28 points, material constraints -30.2 points
  • Claude Opus 4.7: Main up 20.4 points, code execution +41.4 points, material constraints -5.3 points
  • Claude Sonnet 4.6: Main up 18.7 points, code execution +47 points, material constraints -15.8 points
  • Gemini 3.1 Pro: Main down 15.9 points, material constraints -30.8 points
  • GPT-o3: Main up 14.6 points, code execution +22 points, material constraints +5.5 points

Signals to Watch

  • Claude Sonnet 4.6: Material constraints plunge -15.8 points
  • Grok 4: Material constraints plunge -15.8 points
  • Gemini 3.1 Pro: Main plunges -15.9 points
  • Gemini 2.5 Pro: Main plunges -13.1 points
  • DeepSeek V4 Pro: Main plunges -29 points
  • GLM-4.6: Data incomplete (missing multiple evaluation dimensions, API failure/timeout); automatically queued for rerun; not included in this period's ranking

When reading this kind of Smoke brief, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether its integrity rating has moved from pass to warn or fail. Large single-day changes in execution or constraint scores may come from question sampling, or they may be early signals of genuine degradation, and require follow-up run review.


Data source: YZ Index | Run #334 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!