Gemini 2.5 Pro Leads at 87.21: 2026-08-07 Smoke Quick Test Data Briefing

The 2026-08-07 YZ Index Smoke quick test covered 9 models, with Gemini 2.5 Pro ranking first for the day at 87.21 points. Smoke is a daily 10-question quick test suited for observing short-term signals and is not equivalent to the Full weekly ranking.

This Smoke evaluation only covers two main leaderboard dimensions, code execution and material constraints, and the main leaderboard formula is 0.55 × code execution + 0.45 × material constraints. Given the small daily sample size, single-day scores are better treated as monitoring signals rather than long-term conclusions about model capabilities.

Daily Ranking

RankModelMain LeaderboardCode ExecutionMaterial ConstraintsIntegrity
#1Gemini 2.5 Pro87.2194.578.3pass
#2Grok 481.2394.565warn
#3Claude Opus 4.780.9869.595pass
#4GPT-5.574.4894.550pass
#5GPT-o372.3694.545.3pass
#6Claude Sonnet 4.671.4694.543.3pass
#7Gemini 3.1 Pro67.5971.163.3pass
#8Qwen3 Max65.2369.560pass
#9DeepSeek V4 Pro61.2269.551.1pass

Data Interpretation

In today's data, Gemini 2.5 Pro leads the main leaderboard at 87.21, with a balanced combination of 94.5 in code execution and 78.3 in material constraints. Grok 4 scores 94.5 in code execution but 65 in material constraints, yielding a main leaderboard score of 81.23; Claude Opus 4.7 scores 95 in material constraints but 69.5 in code execution, yielding 80.98. GPT-5.5 scores 94.5 in code execution and 50 in material constraints, with a main leaderboard score of 74.48, showing the different emphases of top models across the two dimensions.

Among the notable changes, DeepSeek V4 Pro fell 31 points on the main leaderboard, 30.5 points in code execution, and 31.5 points in material constraints; Claude Sonnet 4.6 fell 20.7 points on the main leaderboard and 39.3 points in material constraints; and Gemini 2.5 Pro rose 15.9 points on the main leaderboard and 23.7 points in code execution. Qwen3 Max fell 15.3 points on the main leaderboard and 18 points in code execution.

Anomalous signals include Grok 4's sharp 17.6-point drop in material constraints, Claude Opus 4.7's sharp 30.5-point drop in code execution, and GPT-5.5's sharp 12.9-point drop on the main leaderboard. These may stem from question sampling fluctuation or genuine regression and require subsequent runs to confirm. Since the Smoke quick test is a small-sample single-day signal, interpretations of changes should remain measured.

Key Changes

  • DeepSeek V4 Pro: main leaderboard down 31 points, code execution -30.5, material constraints -31.5
  • Claude Sonnet 4.6: main leaderboard down 20.7 points, code execution -5.5, material constraints -39.3
  • GPT-o3: main leaderboard down 16.8 points, code execution -5.5, material constraints -30.7
  • Gemini 2.5 Pro: main leaderboard up 15.9 points, code execution +23.7, material constraints +6.4
  • Qwen3 Max: main leaderboard down 15.3 points, code execution -18, material constraints -11.9

Signals to Watch

  • Grok 4: material constraints plunged 17.6 points
  • Claude Opus 4.7: code execution plunged 30.5 points
  • GPT-5.5: main leaderboard plunged 12.9 points
  • GPT-o3: main leaderboard plunged 16.8 points
  • Claude Sonnet 4.6: main leaderboard plunged 20.7 points
  • Qwen3 Max: main leaderboard plunged 15.3 points
  • DeepSeek V4 Pro: main leaderboard plunged 31 points
  • Doubao Pro: data incomplete (multiple evaluation dimensions missing; API failure/timeout), queued for automatic re-run, not ranked this round
  • GLM-4.6: data incomplete (multiple evaluation dimensions missing; API failure/timeout), queued for automatic re-run, not ranked this round

When reading this type of Smoke briefing, the focus should be on two questions: first, whether a model has exposed the same type of weakness on multiple consecutive days; second, whether its integrity rating has moved from pass to warn or fail. Large single-day swings in execution or constraint scores may come from question sampling or may be early signals of genuine regression, requiring subsequent runs to confirm.


Data source: YZ Index | Run #266 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!