The 2026-10-09 YZ Index Smoke quick test covered 14 models, and Claude Sonnet 4.6 ranked first that day with 94.74 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals, and is not equivalent to the conclusions of the Full weekly ranking.
This Smoke evaluation covered only the two main dimensions of Code Execution and Material Constraints, and the Main Index formula is 0.55 × Code Execution + 0.45 × Material Constraints. Due to the small daily sample size, single-day scores are more suitable as monitoring signals rather than long-term conclusions about model capabilities.
Daily Ranking
| Rank | Model | Main Index | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | Claude Sonnet 4.6 | 94.74 | 100 | 88.3 | pass |
| #2 | Claude Opus 4.7 | 86.25 | 75 | 100 | pass |
| #3 | Doubao Pro | 86.25 | 75 | 100 | pass |
| #4 | GPT-6.1 Sol | 85.33 | 83.3 | 87.8 | pass |
| #5 | DeepSeek V4 Pro | 84.9 | 75 | 97 | pass |
| #6 | GPT-6 Astra | 83.98 | 83.3 | 84.8 | pass |
| #7 | Gemini 3.1 Pro | 75 | 75 | 75 | pass |
| #8 | GPT-5.5 | 75 | 75 | 75 | pass |
| #9 | GPT-6 Luna | 74.82 | 58.3 | 95 | pass |
| #10 | Qwen3 Max | 73.63 | 72.5 | 75 | pass |
| #11 | GPT-6 Sol | 72.66 | 75 | 69.8 | pass |
| #12 | GPT-o3 | 71.8 | 58.3 | 88.3 | pass |
| #13 | Grok 4 | 67.24 | 50 | 88.3 | pass |
| #14 | Gemini 2.5 Pro | 65.82 | 58.3 | 75 | pass |
Data Interpretation
In terms of score structure, Claude Sonnet 4.6 leads with a Main Index of 94.74, combining Code Execution 100 with Material Constraints 88.3, showing a clear advantage on Code Execution; Claude Opus 4.7 and Doubao Pro both show a combination of Code Execution 75 and Material Constraints 100, with both having a Main Index of 86.25, showing how high Material Constraints scores can lift rankings. The combination of GPT-6.1 Sol's Code Execution 83.3 and Material Constraints 87.8 gives it a Main Index of 85.33, ranking fourth, while DeepSeek V4 Pro's structure of Code Execution 75 and Material Constraints 97 supports its Main Index of 84.9.
Among notable changes, DeepSeek V4 Pro's Main Index rose 29.2 points and Code Execution rose 50 points; Doubao Pro's Main Index rose 25.6 points, Code Execution rose 25 points, and Material Constraints rose 26.4 points; GPT-6.1 Sol's Main Index rose 24.1 points and Code Execution rose 33.3 points. These increases may stem from daily question-sampling fluctuation or may reflect temporary performance in specific dimensions, and follow-up runs are needed to confirm stability. On anomalous signals, GPT-6 Luna's Main Index plummeted 8.4 points, GPT-o3's Code Execution plummeted 16.7 points, and Grok 4's Code Execution plummeted 20.3 points; these likewise require rerun data to determine whether this is real degradation.
GLM-4.6 did not participate in the ranking due to incomplete data, and overall, the Smoke results remain a single-day, small-sample signal, so interpretation should remain restrained.
Key Changes
- DeepSeek V4 Pro: Main Index up 29.2 points, Code Execution +50 points
- Doubao Pro: Main Index up 25.6 points, Code Execution +25 points, Material Constraints +26.4 points
- GPT-6.1 Sol: Main Index up 24.1 points, Code Execution +33.3 points, Material Constraints +12.8 points
- GPT-6 Astra: Main Index up 22.7 points, Code Execution +33.3 points, Material Constraints +9.8 points
- Gemini 3.1 Pro: Main Index up 16.2 points, Code Execution +25 points, Material Constraints +5.4 points
Signals to Watch
- GPT-6 Luna: Main Index plummeted -8.4 points
- GLM-4.6: Incomplete data (missing integrity, communication dimensions; API failure/timeout), has entered automatic rerun, and is not participating in this period's ranking
- GPT-o3: Code Execution plummeted -16.7 points
- Grok 4: Code Execution plummeted -20.3 points
When reading this type of Smoke brief, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether its integrity rating moves from pass to warn or fail. Large single-day changes in execution or constraint scores may come from question sampling or may be early signals of real degradation, requiring follow-up runs for verification.
Data from: YZ Index | Run #368 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接