Gemini 2.5 Pro Leads with 89.56: 2026-08-04 Smoke Quick-Test Data Brief

On 2026-08-04, the YZ Index Smoke quick-test covered 9 models, with Gemini 2.5 Pro ranking first for the day at 89.56 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals and not equivalent to the Full weekly ranking conclusions.

This Smoke evaluation covers only two main leaderboard dimensions: code execution and material constraint. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraint. Due to the small daily sample size, single-day scores are more suitable as monitoring signals than as long-term conclusions about model capabilities.

Daily Ranking

RankModelMainCode ExecutionMaterial ConstraintIntegrity
#1Gemini 2.5 Pro89.5610076.8pass
#2Claude Opus 4.789.449780.2pass
#3DeepSeek V4 Pro87.919776.8pass
#4GPT-5.587.199775.2pass
#5Claude Sonnet 4.684.949770.2pass
#6Qwen3 Max84.8392.775.2pass
#7Grok 484.3410065.2pass
#8Gemini 3.1 Pro83.19766.1pass
#9GPT-o380.0788.370pass

Key Changes

  • Qwen3 Max: Main leaderboard +25.9, Code Execution +17.7, Material Constraint +35.9
  • Gemini 3.1 Pro: Main leaderboard +22.5, Code Execution +39.4
  • Claude Sonnet 4.6: Main leaderboard +22, Code Execution +47, Material Constraint -8.6
  • DeepSeek V4 Pro: Main leaderboard +17.7, Code Execution +22, Material Constraint +12.5
  • GPT-5.5: Main leaderboard +12.4, Code Execution +13.7, Material Constraint +10.9

Signals to Watch

  • Doubao Pro: incomplete data (missing five evaluation dimensions; API failure/timeout); auto re-run scheduled; not ranked this round
  • GLM-4.6: incomplete data (missing execution and judgment dimensions; API failure/timeout); auto re-run scheduled; not ranked this round

When reading Smoke briefs like this, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large daily swings in execution or constraint scores may stem from question sampling, or may be early signals of genuine degradation, and need to be rechecked in subsequent runs.


Data source: YZ Index | Run #260 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!