On 2026-08-04, the YZ Index Smoke quick-test covered 9 models, with Gemini 2.5 Pro ranking first for the day at 89.56 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals and not equivalent to the Full weekly ranking conclusions.
This Smoke evaluation covers only two main leaderboard dimensions: code execution and material constraint. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraint. Due to the small daily sample size, single-day scores are more suitable as monitoring signals than as long-term conclusions about model capabilities.
Daily Ranking
| Rank | Model | Main | Code Execution | Material Constraint | Integrity |
|---|---|---|---|---|---|
| #1 | Gemini 2.5 Pro | 89.56 | 100 | 76.8 | pass |
| #2 | Claude Opus 4.7 | 89.44 | 97 | 80.2 | pass |
| #3 | DeepSeek V4 Pro | 87.91 | 97 | 76.8 | pass |
| #4 | GPT-5.5 | 87.19 | 97 | 75.2 | pass |
| #5 | Claude Sonnet 4.6 | 84.94 | 97 | 70.2 | pass |
| #6 | Qwen3 Max | 84.83 | 92.7 | 75.2 | pass |
| #7 | Grok 4 | 84.34 | 100 | 65.2 | pass |
| #8 | Gemini 3.1 Pro | 83.1 | 97 | 66.1 | pass |
| #9 | GPT-o3 | 80.07 | 88.3 | 70 | pass |
Key Changes
- Qwen3 Max: Main leaderboard +25.9, Code Execution +17.7, Material Constraint +35.9
- Gemini 3.1 Pro: Main leaderboard +22.5, Code Execution +39.4
- Claude Sonnet 4.6: Main leaderboard +22, Code Execution +47, Material Constraint -8.6
- DeepSeek V4 Pro: Main leaderboard +17.7, Code Execution +22, Material Constraint +12.5
- GPT-5.5: Main leaderboard +12.4, Code Execution +13.7, Material Constraint +10.9
Signals to Watch
- Doubao Pro: incomplete data (missing five evaluation dimensions; API failure/timeout); auto re-run scheduled; not ranked this round
- GLM-4.6: incomplete data (missing execution and judgment dimensions; API failure/timeout); auto re-run scheduled; not ranked this round
When reading Smoke briefs like this, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large daily swings in execution or constraint scores may stem from question sampling, or may be early signals of genuine degradation, and need to be rechecked in subsequent runs.
Data source: YZ Index | Run #260 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接