GPT-o3 Tops with 95.91 Points: 2026-08-09 Smoke Quick-Test Data Brief

On 2026-08-09, the YZ Index Smoke quick test covered 9 models, with GPT-o3 ranking first on the day at 95.91 points. Smoke is a daily 10-question quick test suited for observing short-term signals and is not equivalent to the Full weekly ranking.

This Smoke evaluation only covers two main ranking dimensions: code execution and material constraint. The main ranking formula is 0.55 × code execution + 0.45 × material constraint. Given the small daily sample size, single-day scores are better used as monitoring signals rather than long-term conclusions about model capabilities.

Daily Ranking

RankModelMain ScoreCode ExecutionMaterial ConstraintIntegrity
#1GPT-o395.9110090.9warn
#2Claude Opus 4.792.8510084.1pass
#3GPT-5.588.7510075pass
#4DeepSeek V4 Pro86.510070pass
#5Qwen3 Max85.829969.7pass
#6Gemini 3.1 Pro84.6610065.9pass
#7Grok 484.2799.365.9pass
#8Claude Sonnet 4.669.567562.9pass
#9Gemini 2.5 Pro63.9569.856.8pass

Data Analysis

Today's YZ Index Smoke quick test shows clear divergence among top models in how code execution and material constraint are combined. GPT-o3 has a main score of 95.91 (code execution 100, material constraint 90.9); Claude Opus 4.7 has a main score of 92.85 (code execution 100, material constraint 84.1); GPT-5.5 has a main score of 88.75 (code execution 100, material constraint 75); DeepSeek V4 Pro has a main score of 86.5 (code execution 100, material constraint 70). These models all achieve 100 or near-100 in code execution, while material constraint ranges from 90.9 to 70, forming the basis for their leading main scores.

Several models saw significant declines: Gemini 2.5 Pro fell 19.4 points on the main score, 25 points on code execution, and 12.5 points on material constraint; Claude Sonnet 4.6 fell 15.7 points on the main score and 34.9 points on material constraint; Gemini 3.1 Pro fell 11.7 points on the main score and 25.9 points on material constraint; Grok 4 fell 11 points on the main score and 23.6 points on material constraint; Qwen3 Max fell 10.2 points on the main score and 21.5 points on material constraint. DeepSeek V4 Pro plunged 19.5 points on material constraint, Qwen3 Max plunged 10.2 points on the main score, Gemini 3.1 Pro plunged 11.7 points on the main score, Grok 4 plunged 11 points on the main score, Claude Sonnet 4.6 plunged 15.7 points on the main score, and Gemini 2.5 Pro plunged 19.4 points on the main score.

The above anomalies may stem from question sampling fluctuation or may reflect genuine regression, requiring confirmation in subsequent runs. GLM-4.6 and Doubao Pro were not ranked due to incomplete data. As a small-sample single-day signal, Smoke-related observations are for reference only.

Key Changes

  • Gemini 2.5 Pro: main score down 19.4 points, code execution -25, material constraint -12.5
  • Claude Sonnet 4.6: main score down 15.7 points, material constraint -34.9
  • Gemini 3.1 Pro: main score down 11.7 points, material constraint -25.9
  • Grok 4: main score down 11 points, material constraint -23.6
  • Qwen3 Max: main score down 10.2 points, material constraint -21.5

Signals to Watch

  • DeepSeek V4 Pro: material constraint plunged -19.5 points
  • Qwen3 Max: main score plunged -10.2 points
  • Gemini 3.1 Pro: main score plunged -11.7 points
  • Grok 4: main score plunged -11 points
  • Claude Sonnet 4.6: main score plunged -15.7 points
  • Gemini 2.5 Pro: main score plunged -19.4 points
  • GLM-4.6: data incomplete (missing judgment and integrity dimensions, API failure/timeout), auto re-run initiated, not ranked this round
  • Doubao Pro: data incomplete (multiple evaluation dimensions missing, API failure/timeout), auto re-run initiated, not ranked this round

When reading Smoke briefings like this, focus on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether the integrity rating has shifted from pass to warn or fail. Large single-day swings in execution or constraint scores may stem from question sampling or may be early signals of genuine regression, requiring verification in subsequent runs.


Data source: YZ Index | Run #270 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!