2026-10-01 YZ Index Smoke quick test covered 15 models, with Doubao Pro taking first place for the day at 95.52. Smoke is a daily 10-question quick test, suited for observing short-term signals and not equivalent to Full weekly leaderboard conclusions.
This Smoke evaluation covers only two main leaderboard dimensions: Code Execution and Material Constraints. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraints. Because the daily sample size is small, a single-day score is better used as a monitoring signal rather than a long-term verdict on model capability.
Daily Rankings
| Rank | Model | Main Leaderboard | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | Doubao Pro | 95.52 | 94.3 | 97 | pass |
| #2 | GPT-6 Sol | 89.56 | 95 | 82.9 | pass |
| #3 | GPT-6 Astra | 88.57 | 95 | 80.7 | pass |
| #4 | GPT-6.1 Sol | 88.57 | 95 | 80.7 | pass |
| #5 | Grok 4 | 88.39 | 100 | 74.2 | pass |
| #6 | Claude Sonnet 4.6 | 87.89 | 95 | 79.2 | pass |
| #7 | Claude Opus 4.7 | 86.25 | 75 | 100 | pass |
| #8 | GLM-4.6 | 84 | 75 | 95 | pass |
| #9 | DeepSeek V4 Pro | 81.25 | 70 | 95 | pass |
| #10 | Gemini 3.1 Pro | 81.25 | 70 | 95 | pass |
| #11 | GPT-o3 | 77.39 | 75 | 80.3 | pass |
| #12 | Gemini 2.5 Pro | 77.16 | 70 | 85.9 | pass |
| #13 | GPT-6 Luna | 77.16 | 70 | 85.9 | pass |
| #14 | Qwen3 Max | 77.16 | 70 | 85.9 | pass |
| #15 | GPT-5.5 | 70.25 | 50 | 95 | pass |
Data Interpretation
Analyzing the score structure, Doubao Pro leads with 95.52 on the main leaderboard, combining 94.3 in Code Execution and 97 in Material Constraints for a balanced profile. GPT-6 Sol has 89.56 on the main leaderboard, with 95 in Code Execution but 82.9 in Material Constraints. Grok 4 scores 100 in Code Execution and 74.2 in Material Constraints, for 88.39 on the main leaderboard. Claude Opus 4.7 scores 100 in Material Constraints and 75 in Code Execution, for 86.25 on the main leaderboard. GLM-4.6 has 84 on the main leaderboard, and its combination of 75 in Code Execution and 95 in Material Constraints also stands out. These leading models each emphasize different strengths and weaknesses between Code Execution and Material Constraints.
Compared with the previous same-basis run, GLM-4.6 rose 31.4 points on the main leaderboard, 53 points in Code Execution, and 5 points in Material Constraints. Doubao Pro rose 25 points on the main leaderboard, 22.3 points in Code Execution, and 28.3 points in Material Constraints. DeepSeek V4 Pro rose 24.2 points on the main leaderboard, 48 points in Code Execution, and fell 5 points in Material Constraints. Grok 4 rose 19.6 points on the main leaderboard and 37 points in Code Execution. Gemini 2.5 Pro rose 13.3 points on the main leaderboard, 20 points in Code Execution, and 5 points in Material Constraints. These changes show pronounced fluctuations in some model metrics in the single-day test.
Anomalous signals include Qwen3 Max plunging 24 points in Code Execution and GPT-5.5 plunging 22 points in Code Execution. Possible explanations are question-sampling variance or genuine regression, but follow-up runs are needed for review. Smoke is a small-sample single-day signal; the related analysis remains restrained and does not draw long-term conclusions.
Key Changes
- GLM-4.6: Main Leaderboard up 31.4 points, Code Execution +53 points, Material Constraints +5 points
- Doubao Pro: Main Leaderboard up 25 points, Code Execution +22.3 points, Material Constraints +28.3 points
- DeepSeek V4 Pro: Main Leaderboard up 24.2 points, Code Execution +48 points, Material Constraints -5 points
- Grok 4: Main Leaderboard up 19.6 points, Code Execution +37 points
- Gemini 2.5 Pro: Main Leaderboard up 13.3 points, Code Execution +20 points, Material Constraints +5 points
Signals to Watch
- Qwen3 Max: Code Execution plunges 24 points
- GPT-5.5: Code Execution plunges 22 points
When reading this kind of Smoke brief, focus should be on two questions: first, whether a model exposes the same type of weakness over multiple consecutive days; second, whether its Integrity rating moves from pass into warn or fail. Large single-day changes in execution or constraint scores may come from question sampling or may be early signals of genuine regression, requiring follow-up runs for review.
Data source: YZ Index | Run #354 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接