The September 26, 2026 YZ Index Smoke quick test covered 10 models, with Doubao Pro ranking first for the day at 92.85. Smoke is a daily 10-question quick test suitable for observing short-term signals, and it is not equivalent to the conclusions of the Full weekly ranking.
This Smoke evaluation covered only two main-ranking dimensions: code execution and material constraints. The main-ranking formula is 0.55 × code execution + 0.45 × material constraints. Because the daily sample size is small, a single-day score is better treated as a monitoring signal rather than a long-term conclusion about model capability.
Daily Ranking
| Rank | Model | Main Ranking | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | Doubao Pro | 92.85 | 98.7 | 85.7 | pass |
| #2 | GPT-5.5 | 91.32 | 100 | 80.7 | pass |
| #3 | Claude Opus 4.7 | 89.47 | 100 | 76.6 | pass |
| #4 | GPT-o3 | 87.22 | 100 | 71.6 | pass |
| #5 | Claude Sonnet 4.6 | 86.1 | 95.5 | 74.6 | pass |
| #6 | Gemini 2.5 Pro | 80.35 | 87.5 | 71.6 | pass |
| #7 | DeepSeek V4 Pro | 77.57 | 75 | 80.7 | pass |
| #8 | Qwen3 Max | 75.45 | 76.8 | 73.8 | pass |
| #9 | Gemini 3.1 Pro | 73.47 | 75 | 71.6 | pass |
| #10 | Grok 4 | 70.41 | 75 | 64.8 | pass |
Data Interpretation
In today’s YZ Index Smoke quick test, Doubao Pro ranked first with a main-ranking score of 92.85. Its combination of 98.7 in code execution and 85.7 in material constraints shows a relatively balanced profile across the two dimensions. GPT-5.5 scored 91.32 on the main ranking, with 100 in code execution and 80.7 in material constraints, likewise maintaining a high position through a strong code execution score. Claude Opus 4.7 scored 89.47, GPT-o3 scored 87.22, and Claude Sonnet 4.6 scored 86.1 on the main ranking. All three posted high code execution scores of 100 or 95.5, while their material constraints scores remained in the ranges of 76.6, 71.6, and 74.6 respectively. This structure allowed them to rank near the top under the main-ranking formula of 0.55 × code execution + 0.45 × material constraints.
Compared with the previous run using the same methodology, Claude Sonnet 4.6 rose by 18.3 points on the main ranking, mainly due to a 45.5-point increase in code execution despite a 14.9-point decline in material constraints. Doubao Pro gained 15.9 points on the main ranking, accompanied by a 32-point rise in code execution. Qwen3 Max increased by 15.2 points on the main ranking and by 28.7 points in code execution. GPT-o3 rose by 14.7 points on the main ranking, with code execution up 50 points and material constraints down 28.4 points. Gemini 2.5 Pro gained 12.6 points on the main ranking, with code execution up 37.5 points and material constraints down 17.9 points. These changes may stem from single-day question sampling fluctuations, or they may reflect temporary performance differences in the material constraints dimension. Subsequent runs are needed to recheck and confirm signal stability.
As a small-sample single-day signal, the above score structure and changes are intended only for same-day observation. A restrained reading helps avoid overinterpreting the models’ overall state.
Key Changes
- Claude Sonnet 4.6: main ranking +18.3 points, code execution +45.5 points, material constraints -14.9 points
- Doubao Pro: main ranking +15.9 points, code execution +32 points
- Qwen3 Max: main ranking +15.2 points, code execution +28.7 points
- GPT-o3: main ranking +14.7 points, code execution +50 points, material constraints -28.4 points
- Gemini 2.5 Pro: main ranking +12.6 points, code execution +37.5 points, material constraints -17.9 points
Signals to Watch
- No publishable anomaly signals were identified in this round.
When reading this type of Smoke brief, the focus should be on two questions: first, whether a model exposes the same type of weakness for multiple consecutive days; second, whether its integrity rating moves from pass to warn or fail. Large single-day changes in execution or constraint scores may come from question sampling, or they may be early signals of real degradation, requiring rechecks in subsequent runs.
Data: YZ Index | Run #339 | View Raw Data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接