On 2026-07-30, the YZ Index Smoke quick test covered 11 models, with Claude Opus 4.7 and GPT-5.5 tying for first at 86.5 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals and not equivalent to the Full weekly ranking conclusions.
This Smoke evaluation only covered two main board dimensions: Code Execution and Material Constraint. The main board formula is 0.55 × Code Execution + 0.45 × Material Constraint. Due to the small daily sample size, single-day scores are better used as monitoring signals rather than long-term conclusions about model capabilities.
Daily Rankings
| Rank | Model | Overall Score | Code Execution | Material Constraint | Integrity |
|---|---|---|---|---|---|
| #1 | Claude Opus 4.7 | 86.5 | 100 | 70 | pass |
| #2 | GPT-5.5 | 86.5 | 100 | 70 | pass |
| #3 | GPT-o3 | 79.91 | 75 | 85.9 | pass |
| #4 | Claude Sonnet 4.6 | 78.86 | 96.9 | 56.8 | pass |
| #5 | Grok 4 | 78.01 | 92 | 60.9 | pass |
| #6 | DeepSeek V4 Pro | 76.85 | 75 | 79.1 | pass |
| #7 | Doubao Pro | 75 | 75 | 75 | pass |
| #8 | Gemini 3.1 Pro | 72.75 | 75 | 70 | pass |
| #9 | Gemini 2.5 Pro | 60.79 | 73.7 | 45 | pass |
| #10 | Qwen3 Max | 60.55 | 54.7 | 67.7 | pass |
| #11 | GLM-4.6 | 39.56 | 50 | 26.8 | pass |
Key Changes
- GPT-5.5: Overall Score up 19.8 points, Code Execution +25 points, Material Constraint +13.5 points
- Grok 4: Overall Score down 11.3 points, Code Execution -5.8 points, Material Constraint -18 points
- Claude Opus 4.7: Overall Score up 9.7 points, Code Execution +25 points, Material Constraint -8.9 points
- Claude Sonnet 4.6: Overall Score up 8.1 points, Code Execution +21.9 points, Material Constraint -8.8 points
- Qwen3 Max: Overall Score down 7.5 points, Code Execution -17.8 points, Material Constraint +5.1 points
Signals to Watch
- Grok 4: Overall Score plummeted -11.3 points
- DeepSeek V4 Pro: Code Execution plummeted -25 points
- Doubao Pro: Code Execution plummeted -22 points
- Gemini 2.5 Pro: Material Constraint plummeted -24.8 points
- Qwen3 Max: Code Execution plummeted -17.8 points
- GLM-4.6: Material Constraint plummeted -40.7 points
When reading such Smoke briefs, the focus should be on two questions: First, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Significant daily fluctuations in execution or constraint scores may be due to item sampling or could be early signals of real degradation, requiring subsequent run verification.
Data source: YZ Index | Run #254 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接