On 2026-09-06, the YZ Index Smoke quick test covered 11 models, with Gemini 2.5 Pro ranking first that day at 88.21 points. Smoke is a daily quick test of 10 questions, suited to observing short-term signals and not equivalent to the conclusions of the Full weekly leaderboard.
This Smoke evaluation only covered the two main leaderboard dimensions of code execution and material constraint, with the main leaderboard formula being 0.55 × code execution + 0.45 × material constraint. As the daily sample size is small, single-day scores are better used as monitoring signals rather than long-term conclusions about model capability.
Daily Ranking
| Rank | Model | Main Leaderboard | Code Execution | Material Constraint | Integrity |
|---|---|---|---|---|---|
| #1 | Gemini 2.5 Pro | 88.21 | 90.1 | 85.9 | pass |
| #2 | Gemini 3.1 Pro | 79.19 | 86.7 | 70 | pass |
| #3 | DeepSeek V4 Pro | 77.44 | 91.7 | 60 | pass |
| #4 | GPT-o3 | 72.75 | 75 | 70 | pass |
| #5 | Grok 4 | 68.66 | 75 | 60.9 | pass |
| #6 | Claude Sonnet 4.6 | 68.32 | 70.3 | 65.9 | pass |
| #7 | GPT-5.5 | 67.75 | 70 | 65 | pass |
| #8 | Doubao Pro | 66.16 | 50 | 85.9 | pass |
| #9 | GLM-4.6 | 63.66 | 70 | 55.9 | pass |
| #10 | Claude Opus 4.7 | 61.25 | 50 | 75 | pass |
| #11 | Qwen3 Max | 57.33 | 54.4 | 60.9 | warn |
Data Interpretation
In today's YZ Index Smoke quick test, Gemini 2.5 Pro ranked first on the main leaderboard with 88.21, pairing a balanced combination of 90.1 in code execution and 85.9 in material constraint. Gemini 3.1 Pro's main leaderboard score of 79.19, code execution of 86.7, and material constraint of 70 indicate relatively outstanding code execution. DeepSeek V4 Pro's main leaderboard score of 77.44, code execution of 91.7, and material constraint of 60 reflect a structural profile of strong code execution but weaker material constraint. GPT-o3's main leaderboard score of 72.75, code execution of 75, and material constraint of 70 are relatively balanced overall.
Compared with the previous run under the same criteria, Gemini 2.5 Pro rose +31.4 points on the main leaderboard, +40.1 in code execution, and +20.7 in material constraint; DeepSeek V4 Pro rose +16.6 on the main leaderboard and +41.7 in code execution, but fell -14.2 in material constraint; Claude Opus 4.7 fell -24.9 on the main leaderboard and -50 in code execution, but rose +5.8 in material constraint; GPT-o3 fell -17.9 on the main leaderboard, -25 in code execution, and -9.2 in material constraint; Doubao Pro fell -17.8 on the main leaderboard and -46 in code execution, but rose +16.7 in material constraint. No anomaly signals were reported; the score changes above may stem from question sampling fluctuation or may reflect true regression, and subsequent runs are needed to confirm signal stability.
As a small-sample, single-day signal, Smoke's interpretation above is described only based on the day's data structure and does not render a judgment on models' long-term performance.
Main Changes
- Gemini 2.5 Pro: main leaderboard +31.4 points, code execution +40.1 points, material constraint +20.7 points
- Claude Opus 4.7: main leaderboard -24.9 points, code execution -50 points, material constraint +5.8 points
- GPT-o3: main leaderboard -17.9 points, code execution -25 points, material constraint -9.2 points
- Doubao Pro: main leaderboard -17.8 points, code execution -46 points, material constraint +16.7 points
- DeepSeek V4 Pro: main leaderboard +16.6 points, code execution +41.7 points, material constraint -14.2 points
Signals Requiring Attention
- No publishable anomaly signal was retained this time.
When reading Smoke briefings of this kind, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether its integrity rating has moved from pass to warn or fail. Large single-day swings in execution or constraint scores may come from question sampling or may be early signals of true regression, requiring review in subsequent runs.
Data source: YZ Index | Run #310 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接