On 2026-08-08, the YZ Index Smoke quick test covered 9 models. Claude Opus 4.7 and GPT-o3 tied at 97.66 points for the top spot. Smoke is a daily 10-question quick test suitable for observing short-term signals and is not equivalent to the conclusions of the Full weekly leaderboard.
This Smoke evaluation only covers two main leaderboard dimensions: code execution and material constraints. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraints. Due to the small daily sample size, single-day scores are better used as monitoring signals rather than long-term conclusions about model capabilities.
Daily Ranking
| Rank | Model | Main Leaderboard | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | Claude Opus 4.7 | 97.66 | 100 | 94.8 | pass |
| #2 | GPT-o3 | 97.66 | 100 | 94.8 | pass |
| #3 | Gemini 3.1 Pro | 96.31 | 100 | 91.8 | pass |
| #4 | Qwen3 Max | 96.04 | 100 | 91.2 | pass |
| #5 | Grok 4 | 95.28 | 100 | 89.5 | pass |
| #6 | DeepSeek V4 Pro | 94.07 | 97.8 | 89.5 | pass |
| #7 | GPT-5.5 | 89.65 | 100 | 77 | pass |
| #8 | Claude Sonnet 4.6 | 85.26 | 75 | 97.8 | pass |
| #9 | Gemini 2.5 Pro | 83.33 | 94.8 | 69.3 | pass |
Main Changes
- DeepSeek V4 Pro: main leaderboard +32.9, code execution +28.3, material constraints +38.4
- Qwen3 Max: main leaderboard +30.8, code execution +30.5, material constraints +31.2
- Gemini 3.1 Pro: main leaderboard +28.7, code execution +28.9, material constraints +28.5
- GPT-o3: main leaderboard +25.3, code execution +5.5, material constraints +49.5
- Claude Opus 4.7: main leaderboard +16.7, code execution +30.5
Signals to Watch
- Claude Sonnet 4.6: code execution dropped sharply by -19.5 points
- Doubao Pro: incomplete data (multiple evaluation dimensions missing due to API failure/timeout); auto re-run initiated, not ranked in this round
- GLM-4.6: incomplete data (multiple evaluation dimensions missing due to API failure/timeout); auto re-run initiated, not ranked in this round
When reading this type of Smoke briefing, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large single-day swings in execution or constraint scores may stem from question sampling, or they may be early signals of genuine regression, requiring review in subsequent runs.
Data source: YZ Index | Run #268 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接