The 2026-09-10 YZ Index Smoke quick test covered 10 models, with Gemini 3.1 Pro taking the top spot for the day with 98.35 points. Smoke is a daily quick test of 10 questions, suited to observing short-term signals; it is not equivalent to the conclusions of the Full weekly leaderboard.
This Smoke evaluation covers only two main-leaderboard dimensions — code execution and material constraints — and the main-leaderboard formula is 0.55 × code execution + 0.45 × material constraints. Given the small daily sample size, single-day scores are best treated as monitoring signals rather than as a long-term verdict on model capability.
Daily Rankings
| Rank | Model | Main Leaderboard | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | Gemini 3.1 Pro | 98.35 | 97 | 100 | pass |
| #2 | Doubao Pro | 94.45 | 91.7 | 97.8 | pass |
| #3 | Grok 4 | 93.79 | 88.7 | 100 | pass |
| #4 | Gemini 2.5 Pro | 88.75 | 100 | 75 | pass |
| #5 | DeepSeek V4 Pro | 86.25 | 75 | 100 | pass |
| #6 | GPT-5.5 | 84.6 | 72 | 100 | pass |
| #7 | GPT-o3 | 78.13 | 69.8 | 88.3 | pass |
| #8 | Claude Opus 4.7 | 73.93 | 52.6 | 100 | pass |
| #9 | Claude Sonnet 4.6 | 73.35 | 72 | 75 | pass |
| #10 | Qwen3 Max | 71.34 | 70.8 | 72 | pass |
Data Interpretation
In today's YZ Index Smoke quick test, Gemini 3.1 Pro ranked first with a main-leaderboard score of 98.35, pairing a code-execution score of 97 with a material-constraints score of 100 in a balanced high-level combination. Doubao Pro followed with a main-leaderboard score of 94.45, code execution of 91.7, and material constraints of 97.8. Grok 4 posted a main-leaderboard score of 93.79, code execution of 88.7, and a perfect 100 on material constraints, likewise showing the advantage of a full score on that dimension. Gemini 2.5 Pro reached 88.75 on the main leaderboard with 100 on code execution but only 75 on material constraints, while DeepSeek V4 Pro — 86.25 on the main leaderboard, 75 on code execution, and 100 on material constraints — exhibited a profile dominated by material constraints.
Compared with the previous run under the same methodology, Gemini 3.1 Pro gained 24.1 points on the main leaderboard, 20.5 on code execution, and 28.6 on material constraints; Gemini 2.5 Pro gained 15.2 on the main leaderboard and 30.5 on code execution; DeepSeek V4 Pro gained 9.8 on the main leaderboard, 5.5 on code execution, and 15 on material constraints. This indicates that leading models improved simultaneously on both code execution and material constraints. Claude Sonnet 4.6 fell 8.5 points on the main leaderboard and 22.5 on code execution while gaining 8.6 on material constraints, and GPT-o3 fell 7.9 on the main leaderboard and 24.7 on code execution while gaining 12.6 on material constraints — the abnormal signals are concentrated in sharp code-execution drops.
GPT-5.5 saw a sharp drop of 22.5 points on code execution, GPT-o3 dropped 24.7 points, Claude Opus 4.7 dropped 16.9 points, and Claude Sonnet 4.6 dropped 8.5 points on the main leaderboard. These may stem from question-sampling fluctuation or genuine degradation and need to be confirmed by subsequent runs. GLM-4.6 was not included in the ranking because its data was incomplete. Since Smoke is a small-sample, single-day signal, the changes above are provided for same-day observation only.
Key Changes
- Gemini 3.1 Pro: main leaderboard +24.1 points, code execution +20.5 points, material constraints +28.6 points
- Gemini 2.5 Pro: main leaderboard +15.2 points, code execution +30.5 points
- DeepSeek V4 Pro: main leaderboard +9.8 points, code execution +5.5 points, material constraints +15 points
- Claude Sonnet 4.6: main leaderboard -8.5 points, code execution -22.5 points, material constraints +8.6 points
- GPT-o3: main leaderboard -7.9 points, code execution -24.7 points, material constraints +12.6 points
Signals to Watch
- GPT-5.5: code execution plunged 22.5 points
- GPT-o3: code execution plunged 24.7 points
- Claude Opus 4.7: code execution plunged 16.9 points
- Claude Sonnet 4.6: main leaderboard plunged 8.5 points
- GLM-4.6: data incomplete (several evaluation dimensions missing due to API failure/timeout); automatic re-run has been initiated; not ranked in this round
When reading this kind of Smoke briefing, the focus should be on two questions: first, whether a given model has exposed the same type of weakness on multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large single-day swings in execution or constraint scores may come from question sampling, or may be early signals of genuine degradation, and need to be rechecked in subsequent runs.
Data source: YZ Index | Run #317 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接