GPT-o3 Leads with 90.64: 2026-09-05 Smoke Quick Test Data Briefing

On 2026-09-05, the YZ Index Smoke quick test covered 11 models, with GPT-o3 taking first place for the day with a score of 90.64. Smoke is a daily 10-question quick test suited to short-term signal observation and is not equivalent to the conclusions of the Full weekly ranking.

This Smoke evaluation covered only the two Overall Score dimensions of code execution and material constraints. The Overall Score formula is 0.55 × Code Execution + 0.45 × Material Constraints. Given the relatively small daily sample size, single-day scores are more appropriate as monitoring signals than as a basis for long-term conclusions about model capability.

Daily Ranking

RankModelOverallCode ExecutionMaterial ConstraintsIntegrity
#1GPT-o390.6410079.2pass
#2Claude Opus 4.786.1410069.2pass
#3Doubao Pro83.949669.2pass
#4Grok 483.6398.765.2warn
#5GPT-5.583.1710062.6pass
#6Claude Sonnet 4.679.1398.755.2pass
#7Qwen3 Max72.397569.2warn
#8Gemini 3.1 Pro67.587558.5pass
#9GLM-4.665.247553.3warn
#10DeepSeek V4 Pro60.895074.2pass
#11Gemini 2.5 Pro56.845065.2pass

Data Interpretation

In today's YZ Index Smoke quick test, the leading models showed clear divergence in how code execution and material constraints were paired. GPT-o3 posted an Overall Score of 90.64, with 100 in code execution and 79.2 in material constraints; Claude Opus 4.7 posted an Overall Score of 86.14, also earning 100 in code execution and 69.2 in material constraints; Doubao Pro posted an Overall Score of 83.94, with 96 in code execution and 69.2 in material constraints. These models maintained high levels on the code execution side while keeping material constraint scores relatively balanced, forming the core structure of today's overall top three. Grok 4 and GPT-5.5, by comparison, supported their overall positions with code execution scores of 98.7 and 100, respectively, but recorded 65.2 and 62.6 in material constraints, pointing to a profile more heavily weighted toward code execution.

As for notable changes, Doubao Pro was up 33.9 points overall from the previous run under the same methodology, with code execution up 46 points and material constraints up 19 points. Claude Sonnet 4.6 rose 27 points overall, with code execution up 48.7 points. GPT-5.5 rose 11.5 points overall, with code execution up 25 points and material constraints down 5 points. GPT-o3 rose 8.2 points overall, with material constraints up 15.1 points. Gemini 2.5 Pro, by contrast, fell 18.6 points overall, with code execution down 25 points and material constraints down 10.8 points. Among the anomaly signals, Qwen3 Max's code execution plummeted 21.7 points, GLM-4.6's material constraints plummeted 21 points, and Gemini 2.5 Pro's overall score plummeted 18.6 points. These single-day fluctuations may stem from question-sampling variance or could reflect genuine degradation, and require confirmation through subsequent runs.

As a small-sample single-day signal, the score patterns and anomalies above are intended for same-day observation only and do not constitute a basis for long-term conclusions.

Key Changes

  • Doubao Pro: Overall score +33.9, code execution +46, material constraints +19
  • Claude Sonnet 4.6: Overall score +27, code execution +48.7
  • Gemini 2.5 Pro: Overall score -18.6, code execution -25, material constraints -10.8
  • GPT-5.5: Overall score +11.5, code execution +25, material constraints -5
  • GPT-o3: Overall score +8.2, material constraints +15.1

Signals to Watch

  • Qwen3 Max: Code execution plummeted 21.7 points
  • GLM-4.6: Material constraints plummeted 21 points
  • Gemini 2.5 Pro: Overall score plummeted 18.6 points

When reading Smoke briefings of this kind, the focus should be on two questions: first, whether a model has been exposed for the same type of weakness on multiple consecutive days; second, whether an integrity rating has moved from pass to warn or fail. A large single-day swing in execution or constraint scores may stem from question sampling or may be an early signal of genuine degradation; follow-up runs are needed for verification.


Data source: YZ Index | Run #309 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!