GPT-o3 Leads with 96.93 Points: 2026-08-18 Smoke Quick-Test Data Brief

On 2026-08-18, the YZ Index Smoke quick test covered 10 models, with GPT-o3 ranking first that day with 96.93 points. Smoke is a daily 10-question quick test suitable for observing short-term signals and is not equivalent to the Full weekly leaderboard conclusions.

This Smoke evaluation only covers two main leaderboard dimensions: code execution and material constraint. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraint. Due to the small daily sample size, single-day scores are more suitable as monitoring signals rather than long-term conclusions about model capabilities.

Daily Ranking

RankModelMain LeaderboardCode ExecutionMaterial ConstraintIntegrity
#1GPT-o396.9398.595pass
#2Claude Opus 4.793.6610085.9pass
#3Grok 491.9596.985.9pass
#4Claude Sonnet 4.691.8110081.8pass
#5GPT-5.590.610079.1pass
#6Qwen3 Max86.9895.376.8pass
#7Doubao Pro82.167590.9pass
#8Gemini 2.5 Pro79.917585.9pass
#9Gemini 3.1 Pro71.1951.795pass
#10DeepSeek V4 Pro70.255095pass

Data Interpretation

Among today's top five models on the main leaderboard, GPT-o3 achieved 96.93 points with a balanced combination of 98.5 in code execution and 95 in material constraint, while Claude Opus 4.7 ranked second with 100 in code execution and 85.9 in material constraint. Claude Sonnet 4.6 also scored 100 in code execution but 81.8 in material constraint, and GPT-5.5 scored 100 in code execution with 79.1 in material constraint, showing that leading models generally maintain high scores in the code execution dimension, while differences in material constraint scores have become the main factor distinguishing main leaderboard rankings. The structures of Doubao Pro (75 in code execution, 90.9 in material constraint), Gemini 3.1 Pro (51.7 in code execution, 95 in material constraint), and DeepSeek V4 Pro (50 in code execution, 95 in material constraint) reflect a relatively stronger material constraint profile.

Gemini 2.5 Pro rose 24.2 points on the main leaderboard and 50 points in code execution; DeepSeek V4 Pro rose 19 points on the main leaderboard, 25 points in code execution, and 11.7 points in material constraint; Claude Sonnet 4.6 rose 8.6 points on the main leaderboard and 25 points in code execution. These changes may stem from daily question sampling fluctuations, or they may reflect genuine performance variations in specific dimensions, requiring subsequent runs under the same criteria for confirmation. Qwen3 Max plunged 8.4 points on the main leaderboard, Doubao Pro plunged 12.6 points, and Gemini 3.1 Pro plunged 12.1 points. Both sampling fluctuation and genuine degradation are possible explanations. This Smoke data represents only a small-sample, single-day signal and should not be used for long-term judgments.

GLM-4.6 was excluded from this round's ranking due to incomplete data caused by an API failure. Overall, the combination of each model's code execution and material constraint strengths directly affects its main leaderboard score, and all anomalous signals require further run verification to rule out chance factors.

Key Changes

  • Gemini 2.5 Pro: main leaderboard up 24.2 points, code execution +50 points, material constraint -7.4 points
  • DeepSeek V4 Pro: main leaderboard up 19 points, code execution +25 points, material constraint +11.7 points
  • Doubao Pro: main leaderboard down 12.6 points, code execution -25 points
  • Gemini 3.1 Pro: main leaderboard down 12.1 points, code execution -23.3 points
  • Claude Sonnet 4.6: main leaderboard up 8.6 points, code execution +25 points, material constraint -11.5 points

Signals to Watch

  • Qwen3 Max: main leaderboard plunged -8.4 points
  • Doubao Pro: main leaderboard plunged -12.6 points
  • Gemini 3.1 Pro: main leaderboard plunged -12.1 points
  • GLM-4.6: incomplete data (missing execution dimension, API failure/timeout), auto-retest initiated, not ranked this round

When reading Smoke briefings like this, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large single-day swings in execution or constraint scores may come from question sampling or may be early signals of genuine degradation, requiring subsequent run reviews.


Data source: YZ Index | Run #283 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!