The 2026-08-22 YZ Index Smoke quick test covered 11 models, with Claude Opus 4.7 ranking first that day with a score of 100. Smoke is a daily 10-question quick test, suitable for observing short-term signals, and is not equivalent to the Full weekly ranking conclusions.
This Smoke evaluation only covers two main ranking dimensions: code execution and material constraint. The main ranking formula is 0.55 × Code Execution + 0.45 × Material Constraint. Due to the small daily sample size, single-day scores are more suitable as monitoring signals rather than long-term conclusions about model capabilities.
Daily Rankings
| Rank | Model | Main Ranking | Code Execution | Material Constraint | Integrity |
|---|---|---|---|---|---|
| #1 | Claude Opus 4.7 | 100 | 100 | 100 | pass |
| #2 | Doubao Pro | 99.01 | 100 | 97.8 | pass |
| #3 | Qwen3 Max | 98.65 | 100 | 97 | pass |
| #4 | Grok 4 | 90.01 | 91.4 | 88.3 | pass |
| #5 | GPT-5.5 | 86.25 | 75 | 100 | pass |
| #6 | Claude Sonnet 4.6 | 84.9 | 75 | 97 | pass |
| #7 | Gemini 2.5 Pro | 83.39 | 69.8 | 100 | pass |
| #8 | GPT-o3 | 79.34 | 72 | 88.3 | pass |
| #9 | Gemini 3.1 Pro | 72.5 | 50 | 100 | pass |
| #10 | GLM-4.6 | 71.51 | 50 | 97.8 | pass |
| #11 | DeepSeek V4 Pro | 67.24 | 50 | 88.3 | pass |
Data Interpretation
In today's YZ Index Smoke quick test, Claude Opus 4.7 ranked first on the main ranking with a score of 100, featuring a balanced structure of 100 in code execution and 100 in material constraint. Doubao Pro, with 100 in code execution and 97.8 in material constraint, formed a combination of high code execution paired with medium material constraint, following closely with a main ranking score of 99.01. Qwen3 Max likewise achieved 100 in code execution and 97 in material constraint, scoring 98.65 on the main ranking. This shows that top models predominantly exhibit a pattern of full or near-full scores in code execution with slightly lower material constraint scores. Grok 4 scored 91.4 in code execution and 88.3 in material constraint, with a main ranking score of 90.01. GPT-5.5, by contrast, showed a structure where material constraint (100) outperformed code execution (75), yielding a main ranking score of 86.25, reflecting how different models have complementary strengths and weaknesses across the two indicators.
Gemini 2.5 Pro's main ranking score rose 65 points compared to the previous run under the same methodology, with code execution up 46.3 points and material constraint up 87.9 points, and integrity changing from fail to pass. GLM-4.6's main ranking score rose 38.8 points, with code execution up 50 points and material constraint up 25 points, and integrity changing from warn to pass. Doubao Pro's main ranking score rose 20.4 points, with code execution up 25 points and material constraint up 14.8 points. These increases are concentrated in notable single-dimension improvements in either material constraint or code execution, and require follow-up runs to confirm whether they are single-day sampling fluctuations.
Gemini 2.5 Pro's integrity rating showed an abnormal fail→pass signal, which may stem from question sampling fluctuations or temporary performance changes in a single test, and may also reflect signs of real degradation. The small-sample single-day Smoke data is insufficient to distinguish between the two, and multiple rounds of review are recommended before making a judgment.
Key Changes
- Gemini 2.5 Pro: main ranking up 65 points, code execution +46.3 points, material constraint +87.9 points, integrity fail→pass
- GLM-4.6: main ranking up 38.8 points, code execution +50 points, material constraint +25 points, integrity warn→pass
- Gemini 3.1 Pro: main ranking up 22.5 points, material constraint +50 points, integrity warn→pass
- Doubao Pro: main ranking up 20.4 points, code execution +25 points, material constraint +14.8 points
- Qwen3 Max: main ranking up 19 points, code execution +25 points, material constraint +11.7 points, integrity warn→pass
Signals Requiring Attention
- Gemini 2.5 Pro: integrity rating downgraded to Fail (fail→pass)
When reading this type of Smoke briefing, the focus should be on two questions: first, whether a model has been exposing the same type of weakness for multiple consecutive days; second, whether the integrity rating has shifted from pass to warn or fail. Large single-day swings in execution or constraint scores may result from question sampling or may be early signals of real degradation, requiring verification in subsequent runs.
Data source: YZ Index | Run #289 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接