The YZ Index Smoke quick test on 2026-09-01 covered 11 models, with GPT-o3 ranking first for the day with 96.98 points. Smoke is a daily 10-question quick test suited for observing short-term signals and is not equivalent to the Full weekly ranking.
This Smoke evaluation covers only two main board dimensions: code execution and material constraints. The main board formula is 0.55 × code execution + 0.45 × material constraints. Due to the small daily sample size, single-day scores are better used as monitoring signals rather than long-term conclusions about model capability.
Daily Ranking
| Rank | Model | Main Board | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | GPT-o3 | 96.98 | 94.5 | 100 | warn |
| #2 | Doubao Pro | 87.74 | 88.5 | 86.8 | pass |
| #3 | Claude Sonnet 4.6 | 86.54 | 94.5 | 76.8 | pass |
| #4 | Grok 4 | 84.99 | 72.7 | 100 | pass |
| #5 | Claude Opus 4.7 | 83.23 | 69.5 | 100 | pass |
| #6 | DeepSeek V4 Pro | 83.23 | 69.5 | 100 | pass |
| #7 | Gemini 2.5 Pro | 83.23 | 69.5 | 100 | pass |
| #8 | Gemini 3.1 Pro | 83.23 | 69.5 | 100 | pass |
| #9 | GPT-5.5 | 83.23 | 69.5 | 100 | pass |
| #10 | Qwen3 Max | 77.56 | 70 | 86.8 | pass |
| #11 | GLM-4.6 | 69.48 | 44.5 | 100 | fail |
Key Changes
- DeepSeek V4 Pro: Main board up 21.6 points, code execution +25, material constraints +17.4
- Gemini 2.5 Pro: Main board up 21.6 points, code execution +25, material constraints +17.4
- Gemini 3.1 Pro: Main board up 17 points, code execution +16.7, material constraints +17.4
- GPT-o3: Main board up 16.1 points, material constraints +35.7, integrity pass→warn
- Claude Sonnet 4.6: Main board up 11.1 points, code execution +25, material constraints -5.8
Signals to Watch
- Grok 4: Code execution plummeted by -21.8 points
- GPT-5.5: Code execution plummeted by -25 points
- GLM-4.6: Integrity rating downgraded to Fail (warn→fail)
When reading these Smoke briefs, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether its integrity rating has moved from pass to warn or fail. Large daily swings in execution or constraint scores may come from question sampling, or may be early signs of real degradation, and require follow-up runs for verification.
Data source: YZ Index | Run #304 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接