Gemini 3.1 Pro Takes the Top Spot with 98.35 Points: 2026-09-10 Smoke Quick Test Data Briefing

The 2026-09-10 YZ Index Smoke quick test covered 10 models, with Gemini 3.1 Pro taking the top spot for the day with 98.35 points. Smoke is a daily quick test of 10 questions, suited to observing short-term signals; it is not equivalent to the conclusions of the Full weekly leaderboard.

This Smoke evaluation covers only two main-leaderboard dimensions — code execution and material constraints — and the main-leaderboard formula is 0.55 × code execution + 0.45 × material constraints. Given the small daily sample size, single-day scores are best treated as monitoring signals rather than as a long-term verdict on model capability.

Daily Rankings

RankModelMain LeaderboardCode ExecutionMaterial ConstraintsIntegrity
#1Gemini 3.1 Pro98.3597100pass
#2Doubao Pro94.4591.797.8pass
#3Grok 493.7988.7100pass
#4Gemini 2.5 Pro88.7510075pass
#5DeepSeek V4 Pro86.2575100pass
#6GPT-5.584.672100pass
#7GPT-o378.1369.888.3pass
#8Claude Opus 4.773.9352.6100pass
#9Claude Sonnet 4.673.357275pass
#10Qwen3 Max71.3470.872pass

Data Interpretation

In today's YZ Index Smoke quick test, Gemini 3.1 Pro ranked first with a main-leaderboard score of 98.35, pairing a code-execution score of 97 with a material-constraints score of 100 in a balanced high-level combination. Doubao Pro followed with a main-leaderboard score of 94.45, code execution of 91.7, and material constraints of 97.8. Grok 4 posted a main-leaderboard score of 93.79, code execution of 88.7, and a perfect 100 on material constraints, likewise showing the advantage of a full score on that dimension. Gemini 2.5 Pro reached 88.75 on the main leaderboard with 100 on code execution but only 75 on material constraints, while DeepSeek V4 Pro — 86.25 on the main leaderboard, 75 on code execution, and 100 on material constraints — exhibited a profile dominated by material constraints.

Compared with the previous run under the same methodology, Gemini 3.1 Pro gained 24.1 points on the main leaderboard, 20.5 on code execution, and 28.6 on material constraints; Gemini 2.5 Pro gained 15.2 on the main leaderboard and 30.5 on code execution; DeepSeek V4 Pro gained 9.8 on the main leaderboard, 5.5 on code execution, and 15 on material constraints. This indicates that leading models improved simultaneously on both code execution and material constraints. Claude Sonnet 4.6 fell 8.5 points on the main leaderboard and 22.5 on code execution while gaining 8.6 on material constraints, and GPT-o3 fell 7.9 on the main leaderboard and 24.7 on code execution while gaining 12.6 on material constraints — the abnormal signals are concentrated in sharp code-execution drops.

GPT-5.5 saw a sharp drop of 22.5 points on code execution, GPT-o3 dropped 24.7 points, Claude Opus 4.7 dropped 16.9 points, and Claude Sonnet 4.6 dropped 8.5 points on the main leaderboard. These may stem from question-sampling fluctuation or genuine degradation and need to be confirmed by subsequent runs. GLM-4.6 was not included in the ranking because its data was incomplete. Since Smoke is a small-sample, single-day signal, the changes above are provided for same-day observation only.

Key Changes

  • Gemini 3.1 Pro: main leaderboard +24.1 points, code execution +20.5 points, material constraints +28.6 points
  • Gemini 2.5 Pro: main leaderboard +15.2 points, code execution +30.5 points
  • DeepSeek V4 Pro: main leaderboard +9.8 points, code execution +5.5 points, material constraints +15 points
  • Claude Sonnet 4.6: main leaderboard -8.5 points, code execution -22.5 points, material constraints +8.6 points
  • GPT-o3: main leaderboard -7.9 points, code execution -24.7 points, material constraints +12.6 points

Signals to Watch

  • GPT-5.5: code execution plunged 22.5 points
  • GPT-o3: code execution plunged 24.7 points
  • Claude Opus 4.7: code execution plunged 16.9 points
  • Claude Sonnet 4.6: main leaderboard plunged 8.5 points
  • GLM-4.6: data incomplete (several evaluation dimensions missing due to API failure/timeout); automatic re-run has been initiated; not ranked in this round

When reading this kind of Smoke briefing, the focus should be on two questions: first, whether a given model has exposed the same type of weakness on multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large single-day swings in execution or constraint scores may come from question sampling, or may be early signals of genuine degradation, and need to be rechecked in subsequent runs.


Data source: YZ Index | Run #317 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!