2026-07-28 YZ Index Smoke Quick Test covered 11 models, with Gemini 3.1 Pro scoring 100 to top the day. Smoke is a daily 10-question quick test, suitable for observing short-term signals, not equivalent to the Full weekly ranking conclusions.
This Smoke evaluation only covers the two main leaderboard dimensions of code execution and material constraint. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraint. Due to the small daily sample size, single-day scores are more suitable as monitoring signals rather than long-term conclusions about model capabilities.
Daily Ranking
| Rank | Model | Main | Code Exec | Material Constraint | Integrity |
|---|---|---|---|---|---|
| #1 | Gemini 3.1 Pro | 100 | 100 | 100 | pass |
| #2 | GPT-o3 | 97.75 | 100 | 95 | pass |
| #3 | GPT-5.5 | 93.66 | 100 | 85.9 | pass |
| #4 | Doubao Pro | 92.07 | 95.8 | 87.5 | pass |
| #5 | Grok 4 | 84.47 | 83.3 | 85.9 | pass |
| #6 | DeepSeek V4 Pro | 84 | 75 | 95 | pass |
| #7 | Claude Sonnet 4.6 | 79.91 | 75 | 85.9 | pass |
| #8 | Gemini 2.5 Pro | 77.6 | 70.8 | 85.9 | pass |
| #9 | GLM-4.6 | 77.6 | 70.8 | 85.9 | pass |
| #10 | Qwen3 Max | 59.28 | 37.5 | 85.9 | pass |
| #11 | Claude Opus 4.7 | 56.5 | 25 | 95 | pass |
Data Interpretation
Gemini 3.1 Pro leads with a main score of 100, code execution of 100, and material constraint of 100, showing a balanced and high-performance profile in both dimensions. GPT-o3 scores 97.75 on the main leaderboard, 100 on code execution, and 95 on material constraint, with code execution outperforming material constraint. GPT-5.5 scores 93.66 on the main leaderboard, 100 on code execution, and 85.9 on material constraint, also primarily supported by code execution. Doubao Pro scores 92.07 on the main leaderboard, 95.8 on code execution, and 87.5 on material constraint, with a relatively balanced structure.
On the anomaly front, Gemini 3.1 Pro's main leaderboard increased by 11 points and material constraint by 20.7 points, indicating an overall boost driven by material constraint. Grok 4's main leaderboard increased by 9.3 points and material constraint by 15.7 points; GPT-5.5's main leaderboard increased by 7.4 points and material constraint by 12.7 points; Doubao Pro's main leaderboard increased by 7.1 points and material constraint by 17.3 points, all reflecting a significant rise in material constraint scores.
Among anomaly signals, DeepSeek V4 Pro's code execution plummeted by 25 points, Claude Sonnet 4.6's code execution plummeted by 22 points, Gemini 2.5 Pro's code execution plummeted by 25 points, Qwen3 Max's code execution plummeted by 22 points, and Claude Opus 4.7's main leaderboard plummeted by 9.5 points. These changes may stem from question sampling fluctuations or genuine degradation and require confirmation in subsequent run reviews. As a small-sample single-day signal, interpretations of Smoke should remain restrained.
Key Changes
- Gemini 3.1 Pro: Main leaderboard up 11 points, material constraint +20.7 points
- Claude Opus 4.7: Main leaderboard down 9.5 points, code execution -22 points, material constraint +5.7 points
- Grok 4: Main leaderboard up 9.3 points, material constraint +15.7 points
- GPT-5.5: Main leaderboard up 7.4 points, material constraint +12.7 points
- Doubao Pro: Main leaderboard up 7.1 points, material constraint +17.3 points
Signals to Watch
- DeepSeek V4 Pro: Code execution plunged -25 points
- Claude Sonnet 4.6: Code execution plunged -22 points
- Gemini 2.5 Pro: Code execution plunged -25 points
- Qwen3 Max: Code execution plunged -22 points
- Claude Opus 4.7: Main leaderboard plunged -9.5 points
When reading such Smoke briefs, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether the integrity rating has changed from pass to warn or fail. Large single-day changes in execution or constraint scores may be due to question sampling or early signals of genuine degradation, requiring confirmation in subsequent run reviews.
Data source: YZ Index | Run #250 | View Raw Data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接