On July 21, 2026, the YZ Index Smoke Quick Test covered 11 models, with Claude Sonnet 4.6 and GPT-o3 tying for first place at 96.27 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals and is not equivalent to Full weekly ranking conclusions.
This Smoke evaluation only covers two main leaderboard dimensions: code execution and material constraints. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraints. Due to the small daily sample size, single-day scores are better used as monitoring signals rather than long-term conclusions about model capabilities.
Daily Ranking
| Rank | Model | Main Leaderboard | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | Claude Sonnet 4.6 | 96.27 | 100 | 91.7 | pass |
| #2 | GPT-o3 | 96.27 | 100 | 91.7 | pass |
| #3 | Doubao Pro | 92.49 | 100 | 83.3 | pass |
| #4 | Grok 4 | 88.8 | 93.3 | 83.3 | pass |
| #5 | DeepSeek V4 Pro | 87.87 | 86.7 | 89.3 | pass |
| #6 | GPT-5.5 | 87.67 | 100 | 72.6 | pass |
| #7 | Gemini 2.5 Pro | 80.21 | 70.8 | 91.7 | pass |
| #8 | Qwen3 Max | 79.42 | 85 | 72.6 | pass |
| #9 | Gemini 3.1 Pro | 75.9 | 78.6 | 72.6 | pass |
| #10 | Claude Opus 4.7 | 73.92 | 75 | 72.6 | pass |
| #11 | GLM-4.6 | 62.83 | 50 | 78.5 | pass |
Data Interpretation
Today's top two models, Claude Sonnet 4.6 and GPT-o3, both scored 96.27. They each achieved 100 in code execution and 91.7 in material constraints, showing a balanced structure in the combination of code execution and material constraints. Doubao Pro scored 100 in code execution and 83.3 in material constraints, with a main leaderboard score of 92.49; Grok 4 scored 93.3 in code execution and 83.3 in material constraints, with a main leaderboard score of 88.8; DeepSeek V4 Pro scored 86.7 in code execution and 89.3 in material constraints, with a main leaderboard score of 87.87, indicating differences in the combination of the two metrics across models.
Compared to the previous comparable run, Claude Sonnet 4.6 increased by 31.9 points in the main leaderboard, 28.1 points in code execution, and 36.5 points in material constraints; GPT-o3 increased by 13.6 points in the main leaderboard and 25 points in code execution; Qwen3 Max increased by 12.1 points in the main leaderboard and 19.4 points in code execution; Gemini 2.5 Pro increased by 10.2 points in the main leaderboard and 20.8 points in code execution. In contrast, Claude Opus 4.7 decreased by 26.1 points in the main leaderboard, 25 points in code execution, and 27.4 points in material constraints.
Gemini 3.1 Pro dropped sharply by 17.8 points in material constraints, Claude Opus 4.7 dropped sharply by 26.1 points in the main leaderboard, and GLM-4.6 dropped sharply by 21.5 points in material constraints. These abnormal signals may stem from question sampling fluctuations or could represent genuine degradation; subsequent runs are needed for verification. Smoke Quick Tests are small-sample single-day signals, so interpretations should be kept restrained.
Key Changes
- Claude Sonnet 4.6: Main leaderboard +31.9, Code Execution +28.1, Material Constraints +36.5
- Claude Opus 4.7: Main leaderboard -26.1, Code Execution -25, Material Constraints -27.4
- GPT-o3: Main leaderboard +13.6, Code Execution +25
- Qwen3 Max: Main leaderboard +12.1, Code Execution +19.4, Integrity warn→pass
- Gemini 2.5 Pro: Main leaderboard +10.2, Code Execution +20.8
Signals to Watch
- Gemini 3.1 Pro: Material Constraints dropped sharply by -17.8 points
- Claude Opus 4.7: Main leaderboard dropped sharply by -26.1 points
- GLM-4.6: Material Constraints dropped sharply by -21.5 points
When reading Smoke briefs like this, the focus should be on two questions: First, whether a model exposes the same type of weakness on multiple consecutive days; second, whether the integrity rating changes from pass to warn or fail. Large single-day fluctuations in execution or constraint scores could be due to question sampling or could be early signals of real degradation; subsequent runs need to be reviewed.
Data Source: YZ Index | Run #240 | View Raw Data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接