GPT-o3 Tops with 86.25 Points: 2026-08-25 Smoke Quick Test Data Brief

On August 25, 2026, the YZ Index Smoke quick test covered 11 models, with GPT-o3 taking the top spot at 86.25 points. Smoke is a daily 10-question quick test designed to observe short-term signals and is not equivalent to the Full weekly ranking conclusions.

This Smoke evaluation covers only two main leaderboard dimensions: code execution and material constraints. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraints. Given the small daily sample size, single-day scores are better used as monitoring signals than as long-term judgments of model capability.

Daily Rankings

RankModelMain RankCode ExecutionMaterial ConstraintsIntegrity
#1GPT-o386.2575100pass
#2GPT-5.583.377593.6pass
#3Grok 483.377593.6pass
#4Doubao Pro81.6791.769.4pass
#5Claude Sonnet 4.680.817587.9pass
#6Gemini 2.5 Pro757575pass
#7Claude Opus 4.772.550100pass
#8Gemini 3.1 Pro70.4466.775pass
#9Qwen3 Max69.985094.4pass
#10DeepSeek V4 Pro56.6941.775pass
#11GLM-4.644.982569.4pass

Data Interpretation

In today's YZ Index Smoke quick test, top models showed clear divergence in how they paired code execution with material constraints. GPT-o3 earned a main rank of 86.25 with code execution at 75 and material constraints at 100. Claude Opus 4.7 achieved a main rank of 72.5 with code execution at 50 and material constraints at 100, illustrating how a perfect material constraints score supports the overall ranking. Doubao Pro scored 91.7 in code execution and 69.4 in material constraints, yielding a main rank of 81.67 and reflecting the boost a high code execution score provides under the 0.55 weight. Gemini 2.5 Pro had 75 in both code execution and material constraints, giving a main rank of 75 — the most balanced profile.

Among models with notable changes, GLM-4.6 fell 42.6 points in main rank, 61.1 points in code execution, and 19.9 points in material constraints; Gemini 3.1 Pro fell 26.5 points in main rank, 27.8 points in code execution, and 25 points in material constraints; DeepSeek V4 Pro fell 23.8 points in main rank and 52.8 points in code execution, while rising 11.7 points in material constraints. Claude Sonnet 4.6 rose 22.6 points in main rank, 30.5 points in code execution, and 12.9 points in material constraints; GPT-5.5 rose 18.7 points in main rank and 30.5 points in code execution. These single-day swings may stem from question sampling variance, or they may reflect genuine short-term performance degradation; follow-up runs under the same methodology are needed to confirm.

The Smoke quick test is a small-sample, single-day signal. The score structure and changes above reflect only the characteristics of that day's data and should not be used for long-term inference.

Key Changes

  • GLM-4.6: Main rank -42.6, code execution -61.1, material constraints -19.9, integrity warn→pass
  • Gemini 3.1 Pro: Main rank -26.5, code execution -27.8, material constraints -25
  • DeepSeek V4 Pro: Main rank -23.8, code execution -52.8, material constraints +11.7
  • Claude Sonnet 4.6: Main rank +22.6, code execution +30.5, material constraints +12.9
  • GPT-5.5: Main rank +18.7, code execution +30.5

Signals to Watch

  • No publishable anomaly signals were retained this time.

When reading Smoke briefs of this kind, focus on two questions: first, whether a model has exposed the same type of weakness on consecutive days; second, whether the integrity rating has shifted from pass to warn or fail. Large single-day changes in execution or constraint scores may come from question sampling or may be early signals of genuine degradation, and require follow-up runs to verify.


Data source: YZ Index | Run #294 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!