On 2026-08-29, the YZ Index Smoke quick test covered 11 models, with Grok 4 topping the day at 96.99 points. Smoke is a daily quick test of 10 questions, suitable for observing short-term signals rather than serving as a conclusion equivalent to the Full weekly ranking.
This Smoke evaluation covers only two main board dimensions: code execution and material constraints. The main board formula is 0.55 × Code Execution + 0.45 × Material Constraints. Since the daily sample size is small, single-day scores are better used as monitoring signals rather than definitive long-term judgments of model capability.
Daily Ranking
| Rank | Model | Main Board | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | Grok 4 | 96.99 | 100 | 93.3 | pass |
| #2 | Gemini 3.1 Pro | 96.6 | 99.3 | 93.3 | pass |
| #3 | Gemini 2.5 Pro | 91.99 | 100 | 82.2 | pass |
| #4 | Qwen3 Max | 91.82 | 90.6 | 93.3 | pass |
| #5 | Doubao Pro | 91.38 | 99.3 | 81.7 | pass |
| #6 | Claude Opus 4.7 | 83.24 | 75 | 93.3 | pass |
| #7 | Claude Sonnet 4.6 | 83.24 | 75 | 93.3 | pass |
| #8 | GLM-4.6 | 83.24 | 75 | 93.3 | pass |
| #9 | GPT-o3 | 83.24 | 75 | 93.3 | pass |
| #10 | DeepSeek V4 Pro | 78.24 | 75 | 82.2 | pass |
| #11 | GPT-5.5 | 61.52 | 50 | 75.6 | pass |
Data Interpretation
Among the top three on today's main board, Grok 4 achieved a main board score of 96.99 with a combination of 100 for code execution and 93.3 for material constraints, while Gemini 3.1 Pro scored 96.6 with 99.3 for code execution and 93.3 for material constraints. Both show a balanced, high-level pairing of code execution and material constraints. Gemini 2.5 Pro scored 100 for code execution but only 82.2 for material constraints, resulting in a main board score of 91.99, reflecting a structure where code execution stands out while material constraints are relatively weaker. Qwen3 Max scored 90.6 for code execution and 93.3 for material constraints, with a main board score of 91.82, demonstrating a stronger material constraints pairing.
Among notable changes, Gemini 3.1 Pro rose by +33.1 on the main board, +44.8 in code execution, and +18.7 in material constraints; GLM-4.6 rose by +22.2 on the main board, +25 in code execution, and +18.7 in material constraints; DeepSeek V4 Pro rose by +19.8 on the main board, +28 in code execution, and +9.8 in material constraints; and Grok 4 rose by +11.5 on the main board and +20.9 in material constraints. These increases may stem from day-to-day question sampling fluctuations and require follow-up run verification. GPT-5.5, meanwhile, showed a clear decline, dropping -24.4 on the main board and -47 in code execution.
Regarding abnormal signals, Claude Opus 4.7 plunged -10.3 on the main board, Claude Sonnet 4.6 plunged -22 in code execution, GPT-o3 plunged -25 in code execution, and GPT-5.5 plunged -24.4 on the main board. Since the Smoke quick test is a small-sample single-day signal, these changes may be caused by question sampling fluctuations or may reflect genuine regression; all require follow-up run verification to determine stability.
Key Changes
- Gemini 3.1 Pro: Main board +33.1, Code Execution +44.8, Material Constraints +18.7
- GPT-5.5: Main board -24.4, Code Execution -47
- GLM-4.6: Main board +22.2, Code Execution +25, Material Constraints +18.7
- DeepSeek V4 Pro: Main board +19.8, Code Execution +28, Material Constraints +9.8
- Grok 4: Main board +11.5, Material Constraints +20.9, Integrity warn → pass
Signals to Watch
- Claude Opus 4.7: Main board plunged -10.3
- Claude Sonnet 4.6: Code Execution plunged -22
- GPT-o3: Code Execution plunged -25
- GPT-5.5: Main board plunged -24.4
When reading this type of Smoke brief, the focus should be on two questions: first, whether a particular model has exposed the same type of weakness for multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large swings in single-day execution or constraint scores may come from question sampling or may be early signals of genuine regression, requiring follow-up run verification.
Data source: YZ Index | Run #299 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接