On 2026-09-02, the YZ Index Smoke quick test covered 11 models, with Doubao Pro ranking first with a score of 95.91. Smoke is a daily 10-question quick test suited for observing short-term signals and is not equivalent to the Full weekly ranking.
This Smoke evaluation only covers two main leaderboard dimensions: code execution and material constraint. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraint. Given the small daily sample size, single-day scores are better used as monitoring signals rather than as definitive long-term conclusions about model capabilities.
Daily Ranking
| Rank | Model | Main Score | Code Execution | Material Constraint | Integrity |
|---|---|---|---|---|---|
| #1 | Doubao Pro | 95.91 | 100 | 90.9 | warn |
| #2 | GPT-o3 | 92.49 | 100 | 83.3 | pass |
| #3 | Claude Sonnet 4.6 | 88.39 | 100 | 74.2 | pass |
| #4 | Gemini 2.5 Pro | 88.39 | 100 | 74.2 | pass |
| #5 | Grok 4 | 88.39 | 100 | 74.2 | pass |
| #6 | Gemini 3.1 Pro | 87.67 | 100 | 72.6 | pass |
| #7 | Claude Opus 4.7 | 72.75 | 75 | 70 | pass |
| #8 | GPT-5.5 | 72.03 | 75 | 68.4 | pass |
| #9 | Qwen3 Max | 69.83 | 75 | 63.5 | pass |
| #10 | GLM-4.6 | 68.22 | 62.5 | 75.2 | warn |
| #11 | DeepSeek V4 Pro | 67.49 | 75 | 58.3 | pass |
Data Analysis
In today's YZ Index Smoke quick test, leading models showed clear divergence in how they pair code execution with material constraint. Doubao Pro leads with a main score of 95.91, with code execution at 100 and material constraint at 90.9; GPT-o3 has a main score of 92.49, also scoring 100 on code execution but 83.3 on material constraint; Claude Sonnet 4.6, Gemini 2.5 Pro, and Grok 4 all share a main score of 88.39, with code execution at 100 and material constraint at 74.2. These models rely on perfect code execution scores to lift their overall results, while differences in material constraint scores directly determine the main leaderboard ordering. GLM-4.6, by contrast, pairs code execution at 62.5 with material constraint at 75.2, yielding a main score of 68.22 and demonstrating an alternative path of low code execution combined with high material constraint.
Multiple models showed notable like-for-like changes. DeepSeek V4 Pro's main score fell 15.7 points with material constraint down 41.7 points; GPT-5.5's main score fell 11.2 points with material constraint down 31.6 points; Claude Opus 4.7's main score fell 10.5 points with material constraint down 30 points; Doubao Pro's main score rose 8.2 points with code execution up 11.5 points, while its integrity rating moved from pass to warn. These fluctuations may stem from single-day sampling variance in a small sample, or they may reflect genuine performance degradation, requiring subsequent run reviews for confirmation. Qwen3 Max's main score fell 7.7 points with material constraint down 23.3 points, also needing further data validation.
Overall, models with a code execution score of 100 occupy the top six in the main leaderboard, but the gap in material constraint scores ranging from 90.9 to 72.6 has already created a spread of more than 8 points in the main ranking. As a small-sample signal, the Smoke quick test results are for reference on the day only and do not constitute a basis for long-term judgment.
Key Changes
- DeepSeek V4 Pro: main score down 15.7 points, code execution +5.5 points, material constraint -41.7 points
- GPT-5.5: main score down 11.2 points, code execution +5.5 points, material constraint -31.6 points
- Claude Opus 4.7: main score down 10.5 points, code execution +5.5 points, material constraint -30 points
- Doubao Pro: main score up 8.2 points, code execution +11.5 points, integrity pass→warn
- Qwen3 Max: main score down 7.7 points, code execution +5 points, material constraint -23.3 points
Signals to Watch
- No publishable anomaly signals were retained in this run.
When reading Smoke briefings like this, the focus should be on two questions: first, whether a model has exposed the same type of weakness on consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large single-day swings in execution or constraint scores may stem from question sampling or may be early signals of genuine degradation, requiring review in subsequent runs.
Data source: YZ Index | Run #305 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接