On 2026-08-31, the YZ Index Smoke quick test covered 11 models, with Grok 4 ranking first that day at 89.15 points. Smoke is a daily 10-question quick test suitable for observing short-term signals and is not equivalent to the Full weekly ranking conclusions.
This Smoke evaluation covers only two main ranking dimensions: Code Execution and Material Constraints. The main score formula is 0.55 × Code Execution + 0.45 × Material Constraints. Due to the small daily sample size, single-day scores are more suitable as monitoring signals rather than long-term conclusions about model capabilities.
Daily Ranking
| Rank | Model | Main Score | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | Grok 4 | 89.15 | 94.5 | 82.6 | pass |
| #2 | Doubao Pro | 86.9 | 94.5 | 77.6 | pass |
| #3 | GPT-5.5 | 86.18 | 94.5 | 76 | pass |
| #4 | GPT-o3 | 80.91 | 94.5 | 64.3 | pass |
| #5 | Claude Opus 4.7 | 75.4 | 69.5 | 82.6 | pass |
| #6 | Claude Sonnet 4.6 | 75.4 | 69.5 | 82.6 | pass |
| #7 | Qwen3 Max | 75.4 | 69.5 | 82.6 | pass |
| #8 | Gemini 3.1 Pro | 66.21 | 52.8 | 82.6 | pass |
| #9 | DeepSeek V4 Pro | 61.65 | 44.5 | 82.6 | pass |
| #10 | Gemini 2.5 Pro | 61.65 | 44.5 | 82.6 | pass |
| #11 | GLM-4.6 | 61.65 | 44.5 | 82.6 | warn |
Data Interpretation
The top three models on today's main score show a clear divergence in their combination of Code Execution and Material Constraints. Grok 4 achieved a main score of 89.15 with Code Execution at 94.5 and Material Constraints at 82.6. Doubao Pro and GPT-5.5 both have Code Execution of 94.5, but Material Constraints of 77.6 and 76 respectively, with main scores of 86.9 and 86.18. Claude Opus 4.7, Claude Sonnet 4.6, and Qwen3 Max all share Code Execution of 69.5 and Material Constraints of 82.6, with the same main score of 75.4. Gemini 3.1 Pro has Code Execution of 52.8 and Material Constraints of 82.6, with a main score of 66.21; DeepSeek V4 Pro, Gemini 2.5 Pro, and GLM-4.6 all have Code Execution of 44.5 and Material Constraints of 82.6, with main scores of 61.65, 61.65, and 61.65.
Compared with the previous same-caliber run, GLM-4.6's main score fell 26.9 points and Code Execution dropped 47 points, with Integrity changing from pass to warn. DeepSeek V4 Pro's main score fell 26.4 points and Code Execution dropped 47 points. Gemini 2.5 Pro's main score fell 24.1 points, Code Execution dropped 50 points, and Material Constraints rose 7.6 points. GPT-5.5's main score rose 23.1 points, Code Execution increased 28 points, and Material Constraints increased 17.1 points. Gemini 3.1 Pro's main score fell 19.6 points and Code Execution dropped 38.7 points. Smoke is a small-sample single-day signal; the above score changes may stem from question sampling fluctuations or may reflect genuine degradation, requiring subsequent runs to confirm stability.
Overall, the three models with Code Execution of 94.5 lead the main score, followed by the three models with Code Execution of 69.5. Models with Code Execution of 52.8 or below rank lower, and models with Material Constraints of 82.6 share the same main score when at the same Code Execution level. All models have an Integrity record of pass, except GLM-4.6 with warn.
Key Changes
- GLM-4.6: main score -26.9, Code Execution -47, Integrity pass→warn
- DeepSeek V4 Pro: main score -26.4, Code Execution -47
- Gemini 2.5 Pro: main score -24.1, Code Execution -50, Material Constraints +7.6
- GPT-5.5: main score +23.1, Code Execution +28, Material Constraints +17.1
- Gemini 3.1 Pro: main score -19.6, Code Execution -38.7
Signals to Watch
- No publishable anomaly signals were retained this time.
When reading Smoke briefings of this kind, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether the Integrity rating has moved from pass to warn or fail. Large single-day changes in execution or constraint scores may stem from question sampling or may be early signs of genuine degradation, requiring subsequent runs to confirm.
Data source: YZ Index | Run #302 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接