On 2026-08-27, the YZ Index Smoke quick test covered 11 models, with Claude Opus 4.7 and GPT-o3 tied for first place at 100 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals and not equivalent to the Full weekly ranking conclusions.
This Smoke evaluation only covers two main leaderboard dimensions: code execution and material constraint. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraint. Given the small daily sample size, single-day scores are better used as monitoring signals rather than long-term conclusions about model capability.
Daily Ranking
| Rank | Model | Main Leaderboard | Code Execution | Material Constraint | Integrity |
|---|---|---|---|---|---|
| #1 | Claude Opus 4.7 | 100 | 100 | 100 | pass |
| #2 | GPT-o3 | 100 | 100 | 100 | pass |
| #3 | GPT-5.5 | 95.01 | 100 | 88.9 | pass |
| #4 | Grok 4 | 93.46 | 96.7 | 89.5 | pass |
| #5 | Claude Sonnet 4.6 | 90.78 | 100 | 79.5 | pass |
| #6 | Doubao Pro | 90.78 | 100 | 79.5 | pass |
| #7 | Gemini 3.1 Pro | 88.75 | 100 | 75 | pass |
| #8 | Gemini 2.5 Pro | 87.59 | 99.2 | 73.4 | pass |
| #9 | GLM-4.6 | 86.25 | 75 | 100 | pass |
| #10 | Qwen3 Max | 84.74 | 92.7 | 75 | pass |
| #11 | DeepSeek V4 Pro | 49.03 | 25 | 78.4 | pass |
Key Changes
- GLM-4.6: Main leaderboard up 38.6 points, Code Execution up 37.5, Material Constraint up 40
- DeepSeek V4 Pro: Main leaderboard down 35 points, Code Execution down 50, Material Constraint down 16.6
- Claude Opus 4.7: Main leaderboard up 20.1 points, Code Execution up 25, Material Constraint up 14.1
- GPT-5.5: Main leaderboard up 11 points, Code Execution up 25, Material Constraint down 6.1
- Gemini 2.5 Pro: Main leaderboard down 8.5 points, Material Constraint down 21.6
Signals to Watch
- Claude Sonnet 4.6: Material Constraint plummeted 20.5 points
- Doubao Pro: Material Constraint plummeted 20.5 points
- Gemini 3.1 Pro: Material Constraint plummeted 20 points
- Gemini 2.5 Pro: Main leaderboard plummeted 8.5 points
- DeepSeek V4 Pro: Main leaderboard plummeted 35 points
When reading this type of Smoke briefing, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether its integrity rating has moved from pass to warn or fail. Large swings in single-day execution or constraint scores may stem from question sampling, or may be early signals of genuine degradation, requiring follow-up runs for verification.
Data source: YZ Index | Run #297 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接