Claude Opus 4.7, GPT-5.5, GPT-o3, and Grok 4 Tie at 86.25: 2026-08-16 Smoke Quick Test Data Briefing

On 2026-08-16, the YZ Index Smoke quick test covered 11 models, with Claude Opus 4.7, GPT-5.5, GPT-o3, and Grok 4 tying for first place at 86.25 points. Smoke is a daily 10-question quick test suitable for observing short-term signals and is not equivalent to the full weekly ranking conclusions.

This Smoke evaluation only covers two main leaderboard dimensions: code execution and material constraints. The main leaderboard formula is 0.55 × code execution + 0.45 × material constraints. Due to the small daily sample size, single-day scores are better used as monitoring signals rather than as long-term conclusions about model capabilities.

Daily Ranking

RankModelMain LeaderboardCode ExecutionMaterial ConstraintsIntegrity
#1Claude Opus 4.786.2575100pass
#2GPT-5.586.2575100pass
#3GPT-o386.2575100pass
#4Grok 486.2575100pass
#5Claude Sonnet 4.6757575pass
#6Qwen3 Max757575pass
#7DeepSeek V4 Pro70.4466.775pass
#8Doubao Pro70.4466.775pass
#9Gemini 3.1 Pro70.4466.775pass
#10Gemini 2.5 Pro54.3837.575pass
#11GLM-4.648.7537.562.5pass

Data Interpretation

In today's YZ Index Smoke quick test, four models—Claude Opus 4.7, GPT-5.5, GPT-o3, and Grok 4—tied on the main leaderboard at 86.25, with all achieving 75 in code execution and 100 in material constraints, demonstrating a structure of perfect material constraint scores paired with relatively uniform code execution scores. Claude Sonnet 4.6 and Qwen3 Max achieved a main leaderboard score of 75 with 75 in both code execution and material constraints. DeepSeek V4 Pro, Doubao Pro, and Gemini 3.1 Pro each obtained a main leaderboard score of 70.44 with 66.7 in code execution and 75 in material constraints. Gemini 2.5 Pro's code execution of 37.5 and material constraints of 75 correspond to a main leaderboard score of 54.38, while GLM-4.6's code execution of 37.5 and material constraints of 62.5 correspond to 48.75.

Compared with the previous run under the same methodology, Gemini 3.1 Pro rose 28.9 points on the main leaderboard, 41.7 points in code execution, and 13.3 points in material constraints. Grok 4 rose 27.3 points on the main leaderboard, 25 points in code execution, and 30 points in material constraints. Gemini 2.5 Pro rose 23.1 points on the main leaderboard, 12.5 points in code execution, and 36.1 points in material constraints. GPT-5.5 rose 22.2 points on the main leaderboard and 49.4 points in material constraints. GPT-o3 rose 17.2 points on the main leaderboard and 38.3 points in material constraints. Claude Opus 4.7 recorded a -24.6 point change in code execution, and Claude Sonnet 4.6 recorded a -25 point change in code execution.

The notable decline in the code execution dimension may stem from sampling fluctuation due to the small single-day sample size, or it may reflect genuine performance differences of models under specific constraint scenarios, requiring confirmation through subsequent same-methodology runs. As a daily single-day signal, the current Smoke quick test data is provided solely for observing the day's structure and does not constitute a long-term trend judgment.

Key Changes

  • Gemini 3.1 Pro: main leaderboard up 28.9 points, code execution +41.7, material constraints +13.3
  • Grok 4: main leaderboard up 27.3 points, code execution +25, material constraints +30
  • Gemini 2.5 Pro: main leaderboard up 23.1 points, code execution +12.5, material constraints +36.1
  • GPT-5.5: main leaderboard up 22.2 points, material constraints +49.4
  • GPT-o3: main leaderboard up 17.2 points, material constraints +38.3

Signals to Watch

  • Claude Opus 4.7: code execution plunged -24.6 points
  • Claude Sonnet 4.6: code execution plunged -25 points

When reading Smoke briefings like this, the focus should be on two questions: first, whether a model has exposed the same type of weakness on multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large single-day swings in execution or constraint scores may stem from question sampling, or they may be early signals of genuine degradation, requiring review in subsequent runs.


Data source: YZ Index | Run #280 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!