On 2026-08-19, the YZ Index Smoke quick test covered 10 models, with Claude Opus 4.7 and GPT-5.5 tying for the top spot of the day at 98.35 points. Smoke is a daily 10-question quick test designed for observing short-term signals, and is not equivalent to the Full weekly ranking conclusions.
This Smoke evaluation only covers two main leaderboard dimensions: code execution and material constraint. The main leaderboard formula is 0.55 × code execution + 0.45 × material constraint. Due to the small daily sample size, single-day scores are better used as monitoring signals rather than as long-term conclusions about model capabilities.
Daily Ranking
| Rank | Model | Main Leaderboard | Code Execution | Material Constraint | Integrity |
|---|---|---|---|---|---|
| #1 | Claude Opus 4.7 | 98.35 | 97 | 100 | pass |
| #2 | GPT-5.5 | 98.35 | 97 | 100 | pass |
| #3 | Claude Sonnet 4.6 | 97.53 | 95.5 | 100 | pass |
| #4 | Grok 4 | 95.16 | 91.2 | 100 | pass |
| #5 | Qwen3 Max | 88.75 | 100 | 75 | pass |
| #6 | GPT-o3 | 87.93 | 98.5 | 75 | pass |
| #7 | Doubao Pro | 77.5 | 100 | 50 | pass |
| #8 | Gemini 2.5 Pro | 74.18 | 73.5 | 75 | pass |
| #9 | DeepSeek V4 Pro | 58.78 | 45.5 | 75 | pass |
| #10 | Gemini 3.1 Pro | 48.35 | 47 | 50 | pass |
Data Interpretation
Among the top three models on today's main leaderboard, Claude Opus 4.7 and GPT-5.5 both scored 98.35 with code execution 97 and material constraint 100, while Claude Sonnet 4.6 scored 97.53 with code execution 95.5 and material constraint 100, showing that a perfect material constraint score plays a prominent role in supporting overall ranking. Qwen3 Max has a perfect 100 in code execution but only 75 in material constraint; Doubao Pro has 100 in code execution but 50 in material constraint, reflecting a combination of strong code execution and weaker material constraint. Gemini 2.5 Pro, with code execution 73.5 and material constraint 75, sits in the middle of the leaderboard with 74.18 points.
Gemini 3.1 Pro dropped 22.8 points on the main leaderboard and 45 points on material constraint; DeepSeek V4 Pro dropped 11.5 points on the main leaderboard and 20 points on material constraint; GPT-o3 dropped 9 points on the main leaderboard and 20 points on material constraint; Doubao Pro dropped 40.9 points on material constraint. These anomalies may stem from daily question sampling fluctuations, or may reflect temporary instability in the material constraint dimension, requiring confirmation through subsequent runs using the same methodology. GLM-4.6 did not participate in this period's ranking due to incomplete data caused by an API failure.
The Smoke quick test is a small-sample single-day signal; the changes above only reflect that day's performance and do not constitute a basis for long-term judgments of model capability.
Key Changes
- Gemini 3.1 Pro: main leaderboard down 22.8 points, material constraint -45 points
- DeepSeek V4 Pro: main leaderboard down 11.5 points, material constraint -20 points
- GPT-o3: main leaderboard down 9 points, material constraint -20 points
- GPT-5.5: main leaderboard up 7.8 points, material constraint +20.9 points
- Claude Sonnet 4.6: main leaderboard up 5.7 points, material constraint +18.2 points
Signals to Watch
- GPT-o3: main leaderboard plunged -9 points
- Doubao Pro: material constraint plunged -40.9 points
- DeepSeek V4 Pro: main leaderboard plunged -11.5 points
- Gemini 3.1 Pro: main leaderboard plunged -22.8 points
- GLM-4.6: incomplete data (API failure/timeout), automatic re-run initiated, not ranked in this period
When reading this type of Smoke brief, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large single-day changes in execution or constraint scores may arise from question sampling, or may be early signs of genuine degradation, and require follow-up runs to re-check.
Data source: YZ Index | Run #284 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接