Claude Opus 4.7 and GPT-o3 Tie at 97.66: 2026-08-08 Smoke Quick Test Data Brief

On 2026-08-08, the YZ Index Smoke quick test covered 9 models. Claude Opus 4.7 and GPT-o3 tied at 97.66 points for the top spot. Smoke is a daily 10-question quick test suitable for observing short-term signals and is not equivalent to the conclusions of the Full weekly leaderboard.

This Smoke evaluation only covers two main leaderboard dimensions: code execution and material constraints. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraints. Due to the small daily sample size, single-day scores are better used as monitoring signals rather than long-term conclusions about model capabilities.

Daily Ranking

RankModelMain LeaderboardCode ExecutionMaterial ConstraintsIntegrity
#1Claude Opus 4.797.6610094.8pass
#2GPT-o397.6610094.8pass
#3Gemini 3.1 Pro96.3110091.8pass
#4Qwen3 Max96.0410091.2pass
#5Grok 495.2810089.5pass
#6DeepSeek V4 Pro94.0797.889.5pass
#7GPT-5.589.6510077pass
#8Claude Sonnet 4.685.267597.8pass
#9Gemini 2.5 Pro83.3394.869.3pass

Main Changes

  • DeepSeek V4 Pro: main leaderboard +32.9, code execution +28.3, material constraints +38.4
  • Qwen3 Max: main leaderboard +30.8, code execution +30.5, material constraints +31.2
  • Gemini 3.1 Pro: main leaderboard +28.7, code execution +28.9, material constraints +28.5
  • GPT-o3: main leaderboard +25.3, code execution +5.5, material constraints +49.5
  • Claude Opus 4.7: main leaderboard +16.7, code execution +30.5

Signals to Watch

  • Claude Sonnet 4.6: code execution dropped sharply by -19.5 points
  • Doubao Pro: incomplete data (multiple evaluation dimensions missing due to API failure/timeout); auto re-run initiated, not ranked in this round
  • GLM-4.6: incomplete data (multiple evaluation dimensions missing due to API failure/timeout); auto re-run initiated, not ranked in this round

When reading this type of Smoke briefing, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large single-day swings in execution or constraint scores may stem from question sampling, or they may be early signals of genuine regression, requiring review in subsequent runs.


Data source: YZ Index | Run #268 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!