Claude Sonnet 4.6 and GPT-o3 Tie at 96.27: 2026-07-21 Smoke Quick Test Data Brief

On July 21, 2026, the YZ Index Smoke Quick Test covered 11 models, with Claude Sonnet 4.6 and GPT-o3 tying for first place at 96.27 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals and is not equivalent to Full weekly ranking conclusions.

This Smoke evaluation only covers two main leaderboard dimensions: code execution and material constraints. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraints. Due to the small daily sample size, single-day scores are better used as monitoring signals rather than long-term conclusions about model capabilities.

Daily Ranking

RankModelMain LeaderboardCode ExecutionMaterial ConstraintsIntegrity
#1Claude Sonnet 4.696.2710091.7pass
#2GPT-o396.2710091.7pass
#3Doubao Pro92.4910083.3pass
#4Grok 488.893.383.3pass
#5DeepSeek V4 Pro87.8786.789.3pass
#6GPT-5.587.6710072.6pass
#7Gemini 2.5 Pro80.2170.891.7pass
#8Qwen3 Max79.428572.6pass
#9Gemini 3.1 Pro75.978.672.6pass
#10Claude Opus 4.773.927572.6pass
#11GLM-4.662.835078.5pass

Data Interpretation

Today's top two models, Claude Sonnet 4.6 and GPT-o3, both scored 96.27. They each achieved 100 in code execution and 91.7 in material constraints, showing a balanced structure in the combination of code execution and material constraints. Doubao Pro scored 100 in code execution and 83.3 in material constraints, with a main leaderboard score of 92.49; Grok 4 scored 93.3 in code execution and 83.3 in material constraints, with a main leaderboard score of 88.8; DeepSeek V4 Pro scored 86.7 in code execution and 89.3 in material constraints, with a main leaderboard score of 87.87, indicating differences in the combination of the two metrics across models.

Compared to the previous comparable run, Claude Sonnet 4.6 increased by 31.9 points in the main leaderboard, 28.1 points in code execution, and 36.5 points in material constraints; GPT-o3 increased by 13.6 points in the main leaderboard and 25 points in code execution; Qwen3 Max increased by 12.1 points in the main leaderboard and 19.4 points in code execution; Gemini 2.5 Pro increased by 10.2 points in the main leaderboard and 20.8 points in code execution. In contrast, Claude Opus 4.7 decreased by 26.1 points in the main leaderboard, 25 points in code execution, and 27.4 points in material constraints.

Gemini 3.1 Pro dropped sharply by 17.8 points in material constraints, Claude Opus 4.7 dropped sharply by 26.1 points in the main leaderboard, and GLM-4.6 dropped sharply by 21.5 points in material constraints. These abnormal signals may stem from question sampling fluctuations or could represent genuine degradation; subsequent runs are needed for verification. Smoke Quick Tests are small-sample single-day signals, so interpretations should be kept restrained.

Key Changes

  • Claude Sonnet 4.6: Main leaderboard +31.9, Code Execution +28.1, Material Constraints +36.5
  • Claude Opus 4.7: Main leaderboard -26.1, Code Execution -25, Material Constraints -27.4
  • GPT-o3: Main leaderboard +13.6, Code Execution +25
  • Qwen3 Max: Main leaderboard +12.1, Code Execution +19.4, Integrity warn→pass
  • Gemini 2.5 Pro: Main leaderboard +10.2, Code Execution +20.8

Signals to Watch

  • Gemini 3.1 Pro: Material Constraints dropped sharply by -17.8 points
  • Claude Opus 4.7: Main leaderboard dropped sharply by -26.1 points
  • GLM-4.6: Material Constraints dropped sharply by -21.5 points

When reading Smoke briefs like this, the focus should be on two questions: First, whether a model exposes the same type of weakness on multiple consecutive days; second, whether the integrity rating changes from pass to warn or fail. Large single-day fluctuations in execution or constraint scores could be due to question sampling or could be early signals of real degradation; subsequent runs need to be reviewed.


Data Source: YZ Index | Run #240 | View Raw Data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!