GPT-o3 Tops the List with 77.45 Points: 2026-09-17 Smoke Quick-Test Data Brief

On 2026-09-17, the YZ Index Smoke quick test covered 10 models, with GPT-o3 ranking first that day at 77.45 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals and not equivalent to Full weekly ranking conclusions.

This Smoke evaluation covered only two main leaderboard dimensions: code execution and material constraints. The main leaderboard formula is 0.55 × code execution + 0.45 × material constraints. Due to the small daily sample size, single-day scores are better suited as monitoring signals rather than long-term conclusions about model capabilities.

Daily Rankings

RankModelMain LeaderboardCode ExecutionMaterial ConstraintsIntegrity
#1GPT-o377.457284.1pass
#2Gemini 2.5 Pro77.3573.781.8pass
#3Claude Sonnet 4.676.417281.8pass
#4GPT-5.575.269.881.8pass
#5Doubao Pro70.1973.765.9pass
#6Claude Opus 4.766.764790.9pass
#7Gemini 3.1 Pro66.764790.9pass
#8Grok 465.567.363.3pass
#9Qwen3 Max54.662590.9pass
#10DeepSeek V4 Pro47.845045.2pass

Data Interpretation

Looking at today's main leaderboard data, GPT-o3 scored 77.45 with code execution at 72 and material constraints at 84.1, while Gemini 2.5 Pro scored 77.35 with code execution at 73.7 and material constraints at 81.8. Both show a relatively balanced structure across code and materials. Claude Sonnet 4.6 likewise scored 76.41 with code execution at 72 and material constraints at 81.8. The top models maintained high levels on both indicators, with no obvious single weak spot.

Compared with the previous run using the same methodology, Qwen3 Max fell 25.9 points on the main leaderboard, including a 43.8-point drop in code execution; Claude Sonnet 4.6 fell 23.6 points on the main leaderboard, with code execution down 28 points and material constraints down 18.2 points; DeepSeek V4 Pro fell 19.6 points on the main leaderboard, with material constraints down 49.8 points. These changes may stem from question sampling fluctuations, or they may reflect real fluctuations in model performance on specific tasks, and need to be confirmed by subsequent runs.

The Smoke test is a small-sample single-day signal. The current data reflects only the results of today's 10 questions. Claude Opus 4.7 and Gemini 3.1 Pro both scored 66.76, tied with code execution at 47 and material constraints at 90.9. Grok 4 scored 65.5 on the main leaderboard. The overall structural differences still require more samples to observe.

Main Changes

  • Qwen3 Max: Main leaderboard down 25.9 points, code execution -43.8 points
  • Claude Sonnet 4.6: Main leaderboard down 23.6 points, code execution -28 points, material constraints -18.2 points
  • DeepSeek V4 Pro: Main leaderboard down 19.6 points, code execution +5.2 points, material constraints -49.8 points
  • Claude Opus 4.7: Main leaderboard down 17.2 points, code execution -28 points
  • Grok 4: Main leaderboard down 15.9 points, material constraints -31.7 points

Signals to Watch

  • No publishable anomaly signals were retained this time.

When reading this kind of Smoke brief, the focus should be on two questions: first, whether a model exposes the same type of weakness on multiple consecutive days; second, whether the integrity rating moves from pass to warn or fail. Large single-day changes in execution or constraint scores may come from question sampling, or they may be early signals of real degradation, and need to be confirmed by subsequent runs.


Data source: YZ Index | Run #327 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!