On September 4, 2026, the YZ Index Smoke quick test covered 11 models, and Claude Opus 4.7 ranked first that day with 92.49 points. Smoke is a daily quick test of 10 questions, suitable for observing short-term signals, and is not equivalent to the conclusions of the Full weekly ranking.
This Smoke evaluation only covers two main leaderboard dimensions: code execution and material constraints. The main leaderboard formula is 0.55 × code execution + 0.45 × material constraints. Due to the small daily sample size, single-day scores are more suitable as monitoring signals rather than long-term conclusions about model capabilities.
Daily Ranking
| Rank | Model | Main Leaderboard | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | Claude Opus 4.7 | 92.49 | 100 | 83.3 | pass |
| #2 | Grok 4 | 89.38 | 100 | 76.4 | pass |
| #3 | GPT-o3 | 82.47 | 97.5 | 64.1 | pass |
| #4 | Gemini 2.5 Pro | 75.45 | 75 | 76 | pass |
| #5 | Qwen3 Max | 75.15 | 96.7 | 48.8 | pass |
| #6 | GPT-5.5 | 71.67 | 75 | 67.6 | pass |
| #7 | DeepSeek V4 Pro | 67.24 | 50 | 88.3 | pass |
| #8 | GLM-4.6 | 60.94 | 50 | 74.3 | pass |
| #9 | Gemini 3.1 Pro | 60.65 | 75 | 43.1 | pass |
| #10 | Claude Sonnet 4.6 | 52.16 | 50 | 54.8 | pass |
| #11 | Doubao Pro | 50.09 | 50 | 50.2 | pass |
Data Interpretation
The top three models on today's main leaderboard presented different combinations across the code execution and material constraint metrics. Claude Opus 4.7 achieved a main leaderboard score of 92.49 with 100 in code execution and 83.3 in material constraints; Grok 4 also scored 100 in code execution but 76.4 in material constraints, for a main leaderboard score of 89.38; GPT-o3 scored 97.5 in code execution and 64.1 in material constraints, for a main leaderboard score of 82.47. All three received an integrity rating of pass. DeepSeek V4 Pro scored 50 in code execution and 88.3 in material constraints, for a main leaderboard score of 67.24, reflecting a structural profile in which material constraints are relatively stronger. Gemini 2.5 Pro scored 75 in code execution and 76 in material constraints, for a main leaderboard score of 75.45, sitting at a relatively balanced position across the two metrics.
Multiple models showed notable changes on a comparable basis. Claude Opus 4.7: main leaderboard +33.6 points, code execution +25 points, material constraints +44 points; Grok 4: main leaderboard +27.4 points, code execution +25 points, material constraints +30.3 points; Gemini 2.5 Pro: main leaderboard +27.9 points, code execution +25 points, material constraints +31.5 points; DeepSeek V4 Pro: main leaderboard +22.1 points, material constraints +49 points. Doubao Pro: main leaderboard -29 points, code execution -50 points. These single-day swings should be viewed in light of Smoke's small-sample nature; they may stem from question-sampling fluctuation, or may reflect genuine performance changes, and all require confirmation through subsequent runs.
Anomalous signals were concentrated in several models. GPT-5.5's code execution plunged by 25 points, GLM-4.6's integrity rating shifted from fail to Fail, Claude Sonnet 4.6's main leaderboard score plunged by 13.9 points, and Doubao Pro's main leaderboard score plunged by 29 points. Such changes are common in small-sample single-day tests and should not yet be viewed as long-term regression; they serve only as observation signals, and further validation through multiple rounds of repeated testing is recommended.
Key Changes
- Claude Opus 4.7: main leaderboard up 33.6 points, code execution +25 points, material constraints +44 points
- Doubao Pro: main leaderboard down 29 points, code execution -50 points
- Gemini 2.5 Pro: main leaderboard up 27.9 points, code execution +25 points, material constraints +31.5 points
- Grok 4: main leaderboard up 27.4 points, code execution +25 points, material constraints +30.3 points
- DeepSeek V4 Pro: main leaderboard up 22.1 points, material constraints +49 points
Signals to Watch
- GPT-5.5: code execution plunged by 25 points
- GLM-4.6: integrity rating downgraded to Fail (fail→pass)
- Claude Sonnet 4.6: main leaderboard plunged by 13.9 points
- Doubao Pro: main leaderboard plunged by 29 points
When reading Smoke briefs of this kind, the focus should be on two questions: first, whether a given model has exposed the same type of weakness on multiple consecutive days; second, whether the integrity rating has moved from pass into warn or fail. Large single-day swings in execution or constraint scores may come from question sampling, or may be early signals of genuine degradation, and require review in subsequent runs.
Data source: YZ Index | Run #308 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接