On 2026-08-21, the YZ Index Smoke quick test covered 10 models, with Claude Opus 4.7 ranking first for the day at 95.08 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals and not equivalent to the Full weekly ranking conclusion.
This Smoke evaluation only covers two main leaderboard dimensions: code execution and material constraints. The leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraints. Due to the small daily sample size, single-day scores are better used as monitoring signals rather than as long-term conclusions about model capabilities.
Daily Ranking
| Rank | Model | Main Leaderboard | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | Claude Opus 4.7 | 95.08 | 98.5 | 90.9 | pass |
| #2 | Qwen3 Max | 79.64 | 75 | 85.3 | warn |
| #3 | Doubao Pro | 78.6 | 75 | 83 | pass |
| #4 | Grok 4 | 76.36 | 71.9 | 81.8 | pass |
| #5 | Claude Sonnet 4.6 | 72.15 | 76.6 | 66.7 | pass |
| #6 | GPT-5.5 | 70.09 | 58.3 | 84.5 | pass |
| #7 | DeepSeek V4 Pro | 68.88 | 58.3 | 81.8 | pass |
| #8 | GPT-o3 | 62.47 | 56.8 | 69.4 | warn |
| #9 | Gemini 3.1 Pro | 50 | 50 | 50 | warn |
| #10 | Gemini 2.5 Pro | 18.37 | 23.5 | 12.1 | fail |
Data Interpretation
Looking at the score structure, Claude Opus 4.7 posted 95.08 on the main leaderboard, with code execution at 98.5 and material constraints at 90.9, forming a relatively balanced profile and taking first place. Qwen3 Max posted 79.64 on the main leaderboard, with code execution at 75 and material constraints at 85.3, showing comparatively stronger material constraint performance. Doubao Pro posted 78.6 on the main leaderboard, with code execution at 75 and material constraints at 83, a similar structure. Grok 4 posted 76.36 on the main leaderboard, with code execution at 71.9 and material constraints at 81.8, also leaning toward material constraints.
Among the models with notable fluctuations, Gemini 2.5 Pro fell 44.1 points on the main leaderboard, 21.5 points on code execution, and 71.6 points on material constraints, with integrity dropping from pass to fail. Gemini 3.1 Pro fell 29.5 points on the main leaderboard, 17 points on code execution, and 44.8 points on material constraints, with integrity turning to warn. Claude Sonnet 4.6 fell 22.5 points on the main leaderboard, 15.4 points on code execution, and 31.1 points on material constraints. GPT-o3 fell 22.2 points on the main leaderboard, 19.5 points on code execution, and 25.4 points on material constraints, with integrity turning to warn. GPT-5.5 fell 18.2 points on the main leaderboard and 33.7 points on code execution. These changes may stem from question sampling fluctuation or may reflect genuine degradation, and require confirmation in subsequent runs.
The Smoke test is a small-sample single-day signal; the interpretation above remains measured and draws no long-term conclusions.
Major Changes
- Gemini 2.5 Pro: main leaderboard -44.1 points, code execution -21.5 points, material constraints -71.6 points, integrity pass→fail
- Gemini 3.1 Pro: main leaderboard -29.5 points, code execution -17 points, material constraints -44.8 points, integrity pass→warn
- Claude Sonnet 4.6: main leaderboard -22.5 points, code execution -15.4 points, material constraints -31.1 points
- GPT-o3: main leaderboard -22.2 points, code execution -19.5 points, material constraints -25.4 points, integrity pass→warn
- GPT-5.5: main leaderboard -18.2 points, code execution -33.7 points
Signals to Watch
- Gemini 2.5 Pro: Today's integrity rating is fail (based on today's Smoke data).
When reading Smoke briefings like this, the focus should be on two questions: first, whether a model has exposed the same type of weakness across multiple consecutive days; second, whether the integrity rating has shifted from pass to warn or fail. Large day-to-day swings in execution or constraint scores may come from question sampling or may be early signals of genuine degradation, requiring follow-up runs for confirmation.
Data source: YZ Index | Run #288 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接