GPT-o3 Leads with 96.98 Points: 2026-09-01 Smoke Test Data Brief

The YZ Index Smoke quick test on 2026-09-01 covered 11 models, with GPT-o3 ranking first for the day with 96.98 points. Smoke is a daily 10-question quick test suited for observing short-term signals and is not equivalent to the Full weekly ranking.

This Smoke evaluation covers only two main board dimensions: code execution and material constraints. The main board formula is 0.55 × code execution + 0.45 × material constraints. Due to the small daily sample size, single-day scores are better used as monitoring signals rather than long-term conclusions about model capability.

Daily Ranking

RankModelMain BoardCode ExecutionMaterial ConstraintsIntegrity
#1GPT-o396.9894.5100warn
#2Doubao Pro87.7488.586.8pass
#3Claude Sonnet 4.686.5494.576.8pass
#4Grok 484.9972.7100pass
#5Claude Opus 4.783.2369.5100pass
#6DeepSeek V4 Pro83.2369.5100pass
#7Gemini 2.5 Pro83.2369.5100pass
#8Gemini 3.1 Pro83.2369.5100pass
#9GPT-5.583.2369.5100pass
#10Qwen3 Max77.567086.8pass
#11GLM-4.669.4844.5100fail

Key Changes

  • DeepSeek V4 Pro: Main board up 21.6 points, code execution +25, material constraints +17.4
  • Gemini 2.5 Pro: Main board up 21.6 points, code execution +25, material constraints +17.4
  • Gemini 3.1 Pro: Main board up 17 points, code execution +16.7, material constraints +17.4
  • GPT-o3: Main board up 16.1 points, material constraints +35.7, integrity pass→warn
  • Claude Sonnet 4.6: Main board up 11.1 points, code execution +25, material constraints -5.8

Signals to Watch

  • Grok 4: Code execution plummeted by -21.8 points
  • GPT-5.5: Code execution plummeted by -25 points
  • GLM-4.6: Integrity rating downgraded to Fail (warn→fail)

When reading these Smoke briefs, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether its integrity rating has moved from pass to warn or fail. Large daily swings in execution or constraint scores may come from question sampling, or may be early signs of real degradation, and require follow-up runs for verification.


Data source: YZ Index | Run #304 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!