The 2026-08-06 YZ Index Smoke quick test covered 9 models, with Claude Sonnet 4.6 and DeepSeek V4 Pro tying for first place at 92.17 points. Smoke is a daily 10-question quick test suited for observing short-term signals and does not equate to the Full weekly leaderboard conclusions.
This Smoke evaluation covers only two main leaderboard dimensions — code execution and material constraint — with the main leaderboard formula being 0.55 × Code Execution + 0.45 × Material Constraint. Given the small daily sample size, single-day scores are better suited as monitoring signals rather than long-term conclusions about model capabilities.
Daily Ranking
| Rank | Model | Main Leaderboard | Code Execution | Material Constraint | Integrity |
|---|---|---|---|---|---|
| #1 | Claude Sonnet 4.6 | 92.17 | 100 | 82.6 | pass |
| #2 | DeepSeek V4 Pro | 92.17 | 100 | 82.6 | pass |
| #3 | GPT-o3 | 89.2 | 100 | 76 | pass |
| #4 | Claude Opus 4.7 | 87.36 | 100 | 71.9 | pass |
| #5 | GPT-5.5 | 87.36 | 100 | 71.9 | pass |
| #6 | Grok 4 | 82.99 | 83.3 | 82.6 | pass |
| #7 | Qwen3 Max | 80.48 | 87.5 | 71.9 | pass |
| #8 | Gemini 3.1 Pro | 73.61 | 75 | 71.9 | pass |
| #9 | Gemini 2.5 Pro | 71.3 | 70.8 | 71.9 | pass |
Data Analysis
Among today's top five models on the main leaderboard, Claude Sonnet 4.6 and DeepSeek V4 Pro tied at 92.17, both scoring 100 in code execution and 82.6 in material constraint. GPT-o3 scored 89.2 on the main leaderboard with 100 in code execution and 76 in material constraint. Claude Opus 4.7 and GPT-5.5 both scored 87.36 on the main leaderboard, each with 100 in code execution and 71.9 in material constraint. Grok 4 scored 82.99 on the main leaderboard, with 83.3 in code execution and 82.6 in material constraint, showing a relatively stronger material constraint profile. Qwen3 Max scored 80.48 on the main leaderboard, with 87.5 in code execution and 71.9 in material constraint.
Compared with the previous run under the same methodology, Qwen3 Max rose 30.8 points on the main leaderboard, with code execution up 15.6 points and material constraint up 49.3 points. Grok 4 rose 27.5 points on the main leaderboard, with code execution up 8.3 points and material constraint up 50.9 points. Claude Sonnet 4.6 rose 18.1 points on the main leaderboard, with material constraint up 36.4 points. Gemini 3.1 Pro rose 18.1 points on the main leaderboard, with material constraint up 40.2 points. Claude Opus 4.7 rose 15.4 points on the main leaderboard, with code execution up 25 points. These increases all come from a single-day Smoke sample and require subsequent run verification to distinguish question-sampling fluctuations from genuine changes.
Doubao Pro and GLM-4.6 had incomplete data due to API failures or timeouts, with multiple dimensions missing for each. They were not ranked in this round and have been queued for automatic re-runs. The Smoke quick test is a small-sample single-day signal; the observations above only reflect the day's score composition and make no inferences about long-term model performance.
Key Changes
- Qwen3 Max: Main leaderboard up 30.8 points, code execution +15.6, material constraint +49.3
- Grok 4: Main leaderboard up 27.5 points, code execution +8.3, material constraint +50.9
- Claude Sonnet 4.6: Main leaderboard up 18.1 points, material constraint +36.4
- Gemini 3.1 Pro: Main leaderboard up 18.1 points, material constraint +40.2
- Claude Opus 4.7: Main leaderboard up 15.4 points, code execution +25
Signals to Watch
- Doubao Pro: Incomplete data (missing execution, material constraint, judgment, integrity, and communication dimensions due to API failure/timeout), queued for automatic re-run, not ranked in this round
- GLM-4.6: Incomplete data (missing execution and integrity dimensions due to API failure/timeout), queued for automatic re-run, not ranked in this round
When reading this type of Smoke brief, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large single-day swings in execution or constraint scores could stem from question sampling or could be early signs of genuine degradation, requiring subsequent run verification.
Data source: YZ Index | Run #264 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接