On 2026-09-30, the YZ Index Smoke Quick Test covered 10 models, with Qwen3 Max ranking first for the day at 84.51 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals and not equivalent to Full weekly leaderboard conclusions.
This Smoke evaluation covers only two main board dimensions: Code Execution and Material Constraints. The main board formula is 0.55 × Code Execution + 0.45 × Material Constraints. Because the daily sample size is small, single-day scores are better used as monitoring signals rather than long-term conclusions about model capabilities.
Daily Ranking
| Rank | Model | Main Board | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | Qwen3 Max | 84.51 | 94 | 72.9 | pass |
| #2 | Claude Sonnet 4.6 | 79.41 | 97 | 57.9 | pass |
| #3 | Claude Opus 4.7 | 77.85 | 72 | 85 | pass |
| #4 | GPT-o3 | 76.86 | 72 | 82.8 | pass |
| #5 | Gemini 3.1 Pro | 73.17 | 55.3 | 95 | pass |
| #6 | GPT-5.5 | 70.7 | 72 | 69.1 | pass |
| #7 | Doubao Pro | 70.52 | 72 | 68.7 | pass |
| #8 | Grok 4 | 68.81 | 63 | 75.9 | pass |
| #9 | Gemini 2.5 Pro | 63.91 | 50 | 80.9 | pass |
| #10 | DeepSeek V4 Pro | 57.1 | 22 | 100 | pass |
Data Interpretation
Among today's top three on the main board, Qwen3 Max achieved 84.51 with a combination of 94 in Code Execution and 72.9 in Material Constraints; Claude Sonnet 4.6 reached 79.41 on the strength of a high 97 in Code Execution despite 57.9 in Material Constraints; and Claude Opus 4.7 obtained 77.85 through 85 in Material Constraints with 72 in Code Execution, showing that leading models complement each other with different strengths across the two dimensions. Gemini 3.1 Pro used 95 in Material Constraints to support a main board score of 73.17, while DeepSeek V4 Pro had 100 in Material Constraints but only 22 in Code Execution, reflecting the differences in different models' immediate performance in the Smoke test.
Gemini 2.5 Pro rose 22.7 on the main board, 25 in Code Execution, and 19.9 in Material Constraints; Claude Sonnet 4.6 rose 13.1 on the main board, 10 in Code Execution, and 16.9 in Material Constraints; GPT-5.5 fell 9.5 on the main board, fell 28 in Code Execution, and rose 13.1 in Material Constraints. These like-for-like changes should be viewed in light of the small-sample characteristics. Claude Opus 4.7's Code Execution plunged 28, GPT-o3's Code Execution plunged 21.8, Doubao Pro's Code Execution plunged 24.9, and Grok 4's Code Execution plunged 31.5, all of which may stem from question sampling fluctuations or single-day state variation. GLM-4.6 has incomplete data due to an API failure and has been queued for a rerun; it is not included in this period's ranking.
The Smoke Quick Test is a daily small-sample single-day signal. The score structure and movements above all need to be confirmed by subsequent like-for-like runs, to avoid making inferences about models' true capabilities that go beyond today's data.
Major Changes
- Gemini 2.5 Pro: Main board up 22.7 points, Code Execution +25 points, Material Constraints +19.9 points
- Gemini 3.1 Pro: Main board up 18.2 points, Code Execution +5.3 points, Material Constraints +34 points
- Claude Sonnet 4.6: Main board up 13.1 points, Code Execution +10 points, Material Constraints +16.9 points
- DeepSeek V4 Pro: Main board up 11.4 points, Material Constraints +29 points, Integrity warn→pass
- GPT-5.5: Main board down 9.5 points, Code Execution -28 points, Material Constraints +13.1 points
Signals to Watch
- Claude Opus 4.7: Code Execution plunged -28 points
- GPT-o3: Code Execution plunged -21.8 points
- GPT-5.5: Main board plunged -9.5 points
- Doubao Pro: Code Execution plunged -24.9 points
- Grok 4: Code Execution plunged -31.5 points
- GLM-4.6: Incomplete data (missing execution, judgment dimensions; API failure/timeout), has entered an automatic rerun, and is not included in this period's ranking
When reading this kind of Smoke brief, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; and second, whether its Integrity rating has moved from pass to warn or fail. Large single-day changes in execution or constraint scores may come from question sampling, or they may be early signals of genuine degradation, requiring follow-up runs for confirmation.
Data source: YZ Index | Run #345 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接