GPT-o3 Leads at 83.94 Points: 2026-09-03 Smoke Quick Test Data Brief

The 2026-09-03 YZ Index Smoke quick test covered 11 models, with GPT-o3 ranking first that day at 83.94 points. Smoke is a daily 10-question quick test suited to observing short-term signals and does not equate to the Full weekly ranking conclusions.

This Smoke evaluation only covers the two main ranking dimensions of code execution and material constraint. The main ranking formula is 0.55 × Code Execution + 0.45 × Material Constraint. Given the small daily sample size, single-day scores are better used as monitoring signals rather than long-term determinations of model capability.

Daily Ranking

RankModelMain RankingCode ExecutionMaterial ConstraintIntegrity
#1GPT-o383.9410064.3pass
#2Doubao Pro79.1210053.6pass
#3GPT-5.579.1210053.6pass
#4Claude Sonnet 4.666.097555.2pass
#5Gemini 3.1 Pro65.8483.344.5pass
#6Grok 4627546.1pass
#7Claude Opus 4.758.947539.3pass
#8Qwen3 Max56.096051.3pass
#9GLM-4.648.255046.1fail
#10Gemini 2.5 Pro47.535044.5pass
#11DeepSeek V4 Pro45.195039.3pass

Data Analysis

In today's YZ Index Smoke quick test, leading models showed clear divergence in how code execution and material constraint scores were paired. GPT-o3 combined 100 in code execution with 64.3 in material constraint for a main ranking score of 83.94. Doubao Pro and GPT-5.5 both paired 100 in code execution with 53.6 in material constraint, each finishing with a main ranking score of 79.12. Claude Sonnet 4.6 posted 75 in code execution and 55.2 in material constraint for a main ranking score of 66.09, while Gemini 3.1 Pro posted 83.3 in code execution and 44.5 in material constraint for a main ranking score of 65.84. Among these models, those with perfect code execution scores occupied the leading positions under the main ranking formula of 0.55 × Code Execution + 0.45 × Material Constraint, while models with comparatively lower material constraint scores relied on code execution to lift their overall results.

Several models saw notable declines relative to the previous run under the same criteria. Gemini 2.5 Pro's main ranking score fell 40.9 points, with code execution down 50 points and material constraint down 29.7 points. Grok 4's main ranking score fell 26.4 points, with code execution down 25 points and material constraint down 28.1 points. Claude Sonnet 4.6's main ranking score fell 22.3 points, with code execution down 25 points and material constraint down 19 points. DeepSeek V4 Pro and Gemini 3.1 Pro also posted main ranking declines of 22.3 and 21.8 points respectively, reflecting simultaneous drops in both code execution and material constraint.

This Smoke quick test provides a small-sample, single-day signal. The changes above could stem from question sampling fluctuations, but genuine regression is also possible, and subsequent runs will be needed to confirm stability. GLM-4.6, which received an integrity rating of fail, posted a main ranking score of 48.25. Its profile of 50 in code execution and 46.1 in material constraint placed it in the middle of the leaderboard without additionally affecting the overall ordering.

Key Changes

  • Gemini 2.5 Pro: Main ranking down 40.9 points; code execution -50 points; material constraint -29.7 points
  • Grok 4: Main ranking down 26.4 points; code execution -25 points; material constraint -28.1 points
  • Claude Sonnet 4.6: Main ranking down 22.3 points; code execution -25 points; material constraint -19 points
  • DeepSeek V4 Pro: Main ranking down 22.3 points; code execution -25 points; material constraint -19 points
  • Gemini 3.1 Pro: Main ranking down 21.8 points; code execution -16.7 points; material constraint -28.1 points

Signals to Watch

  • GLM-4.6: Integrity rating today is fail (based on that day's Smoke data).

When reading Smoke briefs like this, focus should be on two questions: first, whether a model has exposed the same type of weakness on consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large single-day swings in execution or constraint scores may come from question sampling or may be early signals of genuine regression, requiring review in subsequent runs.


Data source: YZ Index | Run #307 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!