On 2026-09-07, the YZ Index Smoke quick test covered 11 models, with DeepSeek V4 Pro, Gemini 3.1 Pro, and GLM-4.6 tying for the top spot of the day at 83.49 points. Smoke is a daily 10-question quick test suited to observing short-term signals and is not equivalent to the conclusions of the Full weekly leaderboard.
This Smoke evaluation only covers two main leaderboard dimensions: Code Execution and Material Constraint. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraint. Given the small daily sample size, single-day scores are better used as monitoring signals rather than as long-term verdicts on model capabilities.
Daily Ranking
| Rank | Model | Main Leaderboard | Code Execution | Material Constraint | Integrity |
|---|---|---|---|---|---|
| #1 | DeepSeek V4 Pro | 83.49 | 100 | 63.3 | pass |
| #2 | Gemini 3.1 Pro | 83.49 | 100 | 63.3 | pass |
| #3 | GLM-4.6 | 83.49 | 100 | 63.3 | warn |
| #4 | Doubao Pro | 80.85 | 99.3 | 58.3 | pass |
| #5 | Grok 4 | 80.13 | 99.3 | 56.7 | pass |
| #6 | GPT-o3 | 77.5 | 100 | 50 | pass |
| #7 | GPT-5.5 | 72.03 | 75 | 68.4 | pass |
| #8 | Gemini 2.5 Pro | 71.16 | 73.5 | 68.3 | pass |
| #9 | Claude Sonnet 4.6 | 65.24 | 75 | 53.3 | pass |
| #10 | Claude Opus 4.7 | 61.25 | 50 | 75 | pass |
| #11 | Qwen3 Max | 55.99 | 50 | 63.3 | pass |
Key Changes
- GLM-4.6: Main leaderboard up 19.8 points, code execution +30 points, material constraint +7.4 points, integrity pass→warn
- Gemini 2.5 Pro: Main leaderboard down 17.1 points, code execution -16.6 points, material constraint -17.6 points
- Doubao Pro: Main leaderboard up 14.7 points, code execution +49.3 points, material constraint -27.6 points
- Grok 4: Main leaderboard up 11.5 points, code execution +24.3 points
- DeepSeek V4 Pro: Main leaderboard up 6.1 points, code execution +8.3 points
Signals to Watch
- Doubao Pro: Material constraint plunged -27.6 points
- GPT-o3: Material constraint plunged -20 points
- Gemini 2.5 Pro: Main leaderboard plunged -17.1 points
When reading Smoke briefs like this, the focus should be on two questions: first, whether a model has exposed the same type of weakness across multiple consecutive days; second, whether the integrity rating has shifted from pass to warn or fail. Large single-day swings in execution or constraint scores may stem from question sampling, or they may be early signs of genuine degradation that require follow-up review in subsequent runs.
Data source: YZ Index | Run #312 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接