Gemini 2.5 Pro Tops the Leaderboard with 88.21 Points: 2026-09-06 Smoke Quick Test Data Briefing

On 2026-09-06, the YZ Index Smoke quick test covered 11 models, with Gemini 2.5 Pro ranking first that day at 88.21 points. Smoke is a daily quick test of 10 questions, suited to observing short-term signals and not equivalent to the conclusions of the Full weekly leaderboard.

This Smoke evaluation only covered the two main leaderboard dimensions of code execution and material constraint, with the main leaderboard formula being 0.55 × code execution + 0.45 × material constraint. As the daily sample size is small, single-day scores are better used as monitoring signals rather than long-term conclusions about model capability.

Daily Ranking

RankModelMain LeaderboardCode ExecutionMaterial ConstraintIntegrity
#1Gemini 2.5 Pro88.2190.185.9pass
#2Gemini 3.1 Pro79.1986.770pass
#3DeepSeek V4 Pro77.4491.760pass
#4GPT-o372.757570pass
#5Grok 468.667560.9pass
#6Claude Sonnet 4.668.3270.365.9pass
#7GPT-5.567.757065pass
#8Doubao Pro66.165085.9pass
#9GLM-4.663.667055.9pass
#10Claude Opus 4.761.255075pass
#11Qwen3 Max57.3354.460.9warn

Data Interpretation

In today's YZ Index Smoke quick test, Gemini 2.5 Pro ranked first on the main leaderboard with 88.21, pairing a balanced combination of 90.1 in code execution and 85.9 in material constraint. Gemini 3.1 Pro's main leaderboard score of 79.19, code execution of 86.7, and material constraint of 70 indicate relatively outstanding code execution. DeepSeek V4 Pro's main leaderboard score of 77.44, code execution of 91.7, and material constraint of 60 reflect a structural profile of strong code execution but weaker material constraint. GPT-o3's main leaderboard score of 72.75, code execution of 75, and material constraint of 70 are relatively balanced overall.

Compared with the previous run under the same criteria, Gemini 2.5 Pro rose +31.4 points on the main leaderboard, +40.1 in code execution, and +20.7 in material constraint; DeepSeek V4 Pro rose +16.6 on the main leaderboard and +41.7 in code execution, but fell -14.2 in material constraint; Claude Opus 4.7 fell -24.9 on the main leaderboard and -50 in code execution, but rose +5.8 in material constraint; GPT-o3 fell -17.9 on the main leaderboard, -25 in code execution, and -9.2 in material constraint; Doubao Pro fell -17.8 on the main leaderboard and -46 in code execution, but rose +16.7 in material constraint. No anomaly signals were reported; the score changes above may stem from question sampling fluctuation or may reflect true regression, and subsequent runs are needed to confirm signal stability.

As a small-sample, single-day signal, Smoke's interpretation above is described only based on the day's data structure and does not render a judgment on models' long-term performance.

Main Changes

  • Gemini 2.5 Pro: main leaderboard +31.4 points, code execution +40.1 points, material constraint +20.7 points
  • Claude Opus 4.7: main leaderboard -24.9 points, code execution -50 points, material constraint +5.8 points
  • GPT-o3: main leaderboard -17.9 points, code execution -25 points, material constraint -9.2 points
  • Doubao Pro: main leaderboard -17.8 points, code execution -46 points, material constraint +16.7 points
  • DeepSeek V4 Pro: main leaderboard +16.6 points, code execution +41.7 points, material constraint -14.2 points

Signals Requiring Attention

  • No publishable anomaly signal was retained this time.

When reading Smoke briefings of this kind, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether its integrity rating has moved from pass to warn or fail. Large single-day swings in execution or constraint scores may come from question sampling or may be early signals of true regression, requiring review in subsequent runs.


Data source: YZ Index | Run #310 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!