On August 2, 2026, the YZ Index Smoke quick test covered 10 models, with Doubao Pro ranking first for the day at 96.7 points. Smoke is a daily 10-question quick test suitable for observing short-term signals and does not equate to Full weekly ranking conclusions.
This Smoke evaluation only covers two main ranking dimensions: code execution and material constraint. The main ranking formula is 0.55 × Code Execution + 0.45 × Material Constraint. Due to the small daily sample size, single-day scores are better suited as monitoring signals rather than long-term conclusions about model capabilities.
Daily Rankings
| Rank | Model | Main | Code Execution | Material Constraint | Integrity |
|---|---|---|---|---|---|
| #1 | Doubao Pro | 96.7 | 94 | 100 | pass |
| #2 | Gemini 2.5 Pro | 96.1 | 97 | 95 | pass |
| #3 | GPT-o3 | 96.1 | 97 | 95 | pass |
| #4 | Grok 4 | 96.1 | 97 | 95 | pass |
| #5 | Qwen3 Max | 96.1 | 97 | 95 | pass |
| #6 | Claude Opus 4.7 | 94.45 | 94 | 95 | pass |
| #7 | Claude Sonnet 4.6 | 94.45 | 94 | 95 | pass |
| #8 | DeepSeek V4 Pro | 94.45 | 94 | 95 | warn |
| #9 | Gemini 3.1 Pro | 94.45 | 94 | 95 | pass |
| #10 | GPT-5.5 | 94.45 | 94 | 95 | pass |
Data Interpretation
Today's Smoke quick test shows Doubao Pro leading the main ranking with 96.7, pairing 94 in code execution with a perfect 100 in material constraint. Gemini 2.5 Pro, GPT-o3, Grok 4, and Qwen3 Max all scored 96.1 on the main ranking, with a balanced structure of 97 in code execution and 95 in material constraint. Models from Claude Opus 4.7 to GPT-5.5 scored 94.45 on the main ranking, with 94 in code execution and 95 in material constraint, reflecting score distributions across different strength combinations.
Claude Sonnet 4.6 improved 39.4 points on the main ranking versus the previous run under the same methodology, with code execution up 42.4 points and material constraint up 35.8 points. Gemini 2.5 Pro improved 35.3 points on the main ranking, with code execution up 40.2 points and material constraint up 29.3 points. GPT-5.5, Grok 4, and GPT-o3 also posted main ranking gains ranging from 16.8 to 28.5 points. These changes may stem from question sampling fluctuations and require confirmation in subsequent runs.
GLM-4.6 is excluded from this round's ranking due to incomplete data (missing integrity and communication dimensions, API failure/timeout) and has entered automatic re-run. The Smoke quick test is a small-sample single-day signal, so interpretations of the above fluctuations should remain measured.
Key Changes
- Claude Sonnet 4.6: Main ranking +39.4, code execution +42.4, material constraint +35.8
- Gemini 2.5 Pro: Main ranking +35.3, code execution +40.2, material constraint +29.3
- GPT-5.5: Main ranking +28.5, code execution +35.7, material constraint +19.6
- Grok 4: Main ranking +26.5, code execution +22, material constraint +32.1
- GPT-o3: Main ranking +16.8, code execution +15.2, material constraint +18.8
Signals to Watch
- GLM-4.6: Incomplete data (missing integrity and communication dimensions, API failure/timeout), has entered automatic re-run, not ranked this round
When reading Smoke briefings like this, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large daily swings in execution or constraint scores may come from question sampling or may be early signals of genuine degradation, requiring verification in subsequent runs.
Data source: YZ Index | Run #257 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接