On 2026-08-26, the YZ Index Smoke quick test covered 11 models, with GPT-o3 ranking first that day with 97.75 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals, and is not equivalent to the Full weekly ranking conclusion.
This Smoke evaluation only covers two main ranking dimensions: code execution and material constraint. The main ranking formula is 0.55 × code execution + 0.45 × material constraint. Due to the small daily sample size, single-day scores are more suitable as monitoring signals rather than long-term conclusions about model capabilities.
Daily Ranking
| Ranking | Model | Main Ranking | Code Execution | Material Constraint | Integrity |
|---|---|---|---|---|---|
| #1 | GPT-o3 | 97.75 | 100 | 95 | pass |
| #2 | Gemini 2.5 Pro | 96.1 | 97 | 95 | pass |
| #3 | Grok 4 | 95.17 | 95.3 | 95 | pass |
| #4 | Doubao Pro | 86.25 | 75 | 100 | pass |
| #5 | Claude Sonnet 4.6 | 85.54 | 73.7 | 100 | pass |
| #6 | DeepSeek V4 Pro | 84 | 75 | 95 | pass |
| #7 | Gemini 3.1 Pro | 84 | 75 | 95 | pass |
| #8 | GPT-5.5 | 84 | 75 | 95 | pass |
| #9 | Claude Opus 4.7 | 79.91 | 75 | 85.9 | pass |
| #10 | Qwen3 Max | 79.19 | 73.7 | 85.9 | pass |
| #11 | GLM-4.6 | 47.63 | 37.5 | 60 | pass |
Data Interpretation
In today's YZ Index Smoke quick test, top models showed clear differentiation in the combination of code execution and material constraint. GPT-o3 achieved a main ranking of 97.75 with code execution 100 and material constraint 95; Gemini 2.5 Pro's code execution 97 and material constraint 95 corresponded to a main ranking of 96.1; Grok 4's code execution 95.3 and material constraint 95 corresponded to a main ranking of 95.17. All three maintained a relatively high balance across both metrics. In contrast, Doubao Pro had code execution 75 and material constraint 100, with a main ranking of 86.25; Claude Sonnet 4.6 had code execution 73.7 and material constraint 100, with a main ranking of 85.54, showing a structural characteristic of full marks in material constraint but relatively lower code execution.
Multiple models showed significant changes. DeepSeek V4 Pro's main ranking rose 27.3 points, code execution rose 33.3 points, and material constraint rose 20 points; Gemini 2.5 Pro's main ranking rose 21.1 points, code execution rose 22 points, and material constraint rose 20 points; Gemini 3.1 Pro's main ranking rose 13.6 points, code execution rose 8.3 points, and material constraint rose 20 points; Grok 4's main ranking rose 11.8 points and code execution rose 20.3 points; GPT-o3's main ranking rose 11.5 points, code execution rose 25 points, and material constraint fell 5 points. These single-day changes may stem from question sampling fluctuations, or may reflect genuine performance differences of models on specific tasks, requiring subsequent re-verification under the same criteria.
Regarding anomaly signals, Doubao Pro's code execution plummeted 16.7 points, while its material constraint remained at 100, with a main ranking of 86.25. Such small-sample single-day signals should be viewed cautiously to avoid over-interpretation.
Key Changes
- DeepSeek V4 Pro: main ranking +27.3, code execution +33.3, material constraint +20
- Gemini 2.5 Pro: main ranking +21.1, code execution +22, material constraint +20
- Gemini 3.1 Pro: main ranking +13.6, code execution +8.3, material constraint +20
- Grok 4: main ranking +11.8, code execution +20.3
- GPT-o3: main ranking +11.5, code execution +25, material constraint -5
Signals to Watch
- Doubao Pro: code execution plummeted -16.7 points
When reading these Smoke briefings, focus should be on two questions: first, whether a particular model has exposed the same type of weakness for multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large single-day changes in execution or constraint scores may come from question sampling or may be early signals of genuine degradation, requiring subsequent run re-verification.
Data source: YZ Index | Run #295 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接