On 2026-09-09, the YZ Index Smoke quick test covered 10 models, and Doubao Pro ranked first that day with 89.91 points. Smoke is a daily 10-question quick test suited for observing short-term signals and is not equivalent to the Full weekly ranking conclusions.
This Smoke evaluation only covers the two main benchmark dimensions of code execution and material constraints. The main benchmark formula is 0.55 × Code Execution + 0.45 × Material Constraints. Given the small daily sample size, single-day scores are better used as monitoring signals rather than as long-term conclusions about model capabilities.
Daily Ranking
| Rank | Model | Main Score | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | Doubao Pro | 89.91 | 94.5 | 84.3 | pass |
| #2 | Grok 4 | 88.42 | 93.5 | 82.2 | warn |
| #3 | GPT-o3 | 86.04 | 94.5 | 75.7 | pass |
| #4 | GPT-5.5 | 84.11 | 94.5 | 71.4 | pass |
| #5 | Claude Sonnet 4.6 | 81.86 | 94.5 | 66.4 | pass |
| #6 | Claude Opus 4.7 | 80.98 | 69.5 | 95 | pass |
| #7 | DeepSeek V4 Pro | 76.48 | 69.5 | 85 | pass |
| #8 | Gemini 3.1 Pro | 74.21 | 76.5 | 71.4 | pass |
| #9 | Gemini 2.5 Pro | 73.6 | 69.5 | 78.6 | pass |
| #10 | Qwen3 Max | 72.97 | 69.5 | 77.2 | pass |
Main Changes
- GPT-5.5: Main score up 27.7 points; code execution +44.5 points; material constraints +7.1 points
- GPT-o3: Main score up 24.8 points; code execution +44.5 points
- Claude Opus 4.7: Main score up 22.2 points; code execution +44.5 points; material constraints -5 points
- Qwen3 Max: Main score up 19.9 points; code execution +25.7 points; material constraints +12.9 points
- Doubao Pro: Main score up 15.6 points; code execution +20.7 points; material constraints +9.3 points
Signals to Watch
- Grok 4: Material constraints plunged -17.8 points
- Claude Sonnet 4.6: Material constraints plunged -33.6 points
- DeepSeek V4 Pro: Material constraints plunged -15 points
- GLM-4.6: Data incomplete (several evaluation dimensions missing; API failure/timeout); auto-retry initiated; not ranked this round
When reading Smoke briefs like this, the focus should be on two questions: first, whether a given model has exposed the same type of weakness on multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large single-day swings in execution or material-constraint scores may stem from question sampling, or may be early signs of genuine degradation, and require follow-up runs for confirmation.
Data source: YZ Index | Run #315 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接