Gemini 2.5 Pro Leads with 100 Points: 2026-09-15 Smoke Quick Test Data Brief

The 2026-09-15 YZ Index Smoke quick test covered 10 models, with Gemini 2.5 Pro ranking first for the day at 100 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals, and is not equivalent to the conclusions of the Full weekly leaderboard.

This Smoke evaluation covered only two main leaderboard dimensions: code execution and material constraints. The main leaderboard formula is 0.55 × code execution + 0.45 × material constraints. Because the daily sample size is small, single-day scores are better used as monitoring signals rather than long-term conclusions about model capability.

Daily Ranking

RankModelMain LeaderboardCode ExecutionMaterial ConstraintsIntegrity
#1Gemini 2.5 Pro100100100pass
#2GPT-o388.7510075pass
#3Grok 488.7510075pass
#4Claude Opus 4.783.9410064.3pass
#5Doubao Pro83.9410064.3pass
#6GPT-5.583.9410064.3pass
#7Claude Sonnet 4.677.510050pass
#8Qwen3 Max72.6910039.3pass
#9Gemini 3.1 Pro70.197564.3pass
#10DeepSeek V4 Pro61.255075pass

Key Changes

  • Gemini 2.5 Pro: main leaderboard up 29.8 points, code execution +25 points, material constraints +35.7 points
  • DeepSeek V4 Pro: main leaderboard down 29.8 points, code execution -50 points, material constraints -5.2 points
  • GPT-o3: main leaderboard up 18.6 points, code execution +25 points, material constraints +10.7 points
  • Gemini 3.1 Pro: main leaderboard down 16.8 points, code execution -25 points, material constraints -6.8 points
  • Claude Opus 4.7: main leaderboard down 7.2 points, material constraints -15.9 points

Signals to Watch

  • Claude Opus 4.7: material constraints plunged by -15.9 points
  • Doubao Pro: material constraints plunged by -15.9 points
  • GPT-5.5: material constraints plunged by -15.9 points
  • Qwen3 Max: material constraints plunged by -40.9 points
  • Gemini 3.1 Pro: main leaderboard plunged by -16.8 points
  • DeepSeek V4 Pro: main leaderboard plunged by -29.8 points
  • GLM-4.6: incomplete data, with several evaluation dimensions missing due to API failure/timeout; it has entered automatic rerun and is not included in this ranking

When reading this type of Smoke brief, the focus should be on two questions: first, whether a given model exposes the same type of weakness over several consecutive days; second, whether its integrity rating moves from pass to warn or fail. Large single-day changes in execution or constraint scores may come from question sampling, or they may be early signals of real degradation, and require follow-up runs for review.


Data source: YZ Index | Run #324 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!