On 2026-08-30, the YZ Index Smoke quick test covered 11 models, with Claude Opus 4.7 ranking first for the day with 93.08 points. Smoke is a daily 10-question quick test suited for observing short-term signals, and is not equivalent to Full weekly ranking conclusions.
This Smoke evaluation only covers two main leaderboard dimensions: code execution and material constraint. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraint. Due to the small daily sample size, single-day scores are more suitable as monitoring signals rather than long-term conclusions about model capabilities.
Daily Ranking
| Rank | Model | Leaderboard | Code Execution | Material Constraint | Integrity |
|---|---|---|---|---|---|
| #1 | Claude Opus 4.7 | 93.08 | 91.5 | 95 | pass |
| #2 | GLM-4.6 | 88.58 | 91.5 | 85 | pass |
| #3 | Grok 4 | 88.14 | 90.7 | 85 | pass |
| #4 | DeepSeek V4 Pro | 88.08 | 91.5 | 83.9 | pass |
| #5 | Doubao Pro | 85.89 | 90.7 | 80 | pass |
| #6 | Gemini 3.1 Pro | 85.83 | 91.5 | 78.9 | pass |
| #7 | Gemini 2.5 Pro | 85.73 | 94.5 | 75 | pass |
| #8 | Qwen3 Max | 80.07 | 91.5 | 66.1 | pass |
| #9 | Claude Sonnet 4.6 | 70.33 | 66.5 | 75 | pass |
| #10 | GPT-o3 | 69.56 | 66.5 | 73.3 | pass |
| #11 | GPT-5.5 | 63.08 | 66.5 | 58.9 | pass |
Data Interpretation
Among today's top five models on the main leaderboard, Claude Opus 4.7 achieved a leaderboard score of 93.08 with a code execution score of 91.5 and a material constraint score of 95, the highest material constraint score among all models. GLM-4.6 and DeepSeek V4 Pro also scored 91.5 in code execution, but their material constraint scores were 85 and 83.9 respectively, resulting in leaderboard scores of 88.58 and 88.08. Gemini 2.5 Pro scored 94.5 in code execution but only 75 in material constraint, landing at 85.73 on the leaderboard, showing that the combination of strengths in code execution versus material constraint directly affects overall ranking.
GPT-o3 dropped 13.7 points on the leaderboard, with code execution down 8.5 points and material constraint down 20 points. Claude Sonnet 4.6 dropped 12.9 points on the leaderboard, with code execution down 8.5 points and material constraint down 18.3 points. Qwen3 Max dropped 27.2 points in material constraint. Gemini 3.1 Pro dropped 10.8 points on the leaderboard, with code execution down 7.8 points and material constraint down 14.4 points. These changes may stem from question sampling fluctuations or could reflect genuine single-day performance differences, requiring follow-up runs under the same methodology for confirmation.
As for anomalous signals, Grok 4 plunged 8.8 points on the leaderboard, Gemini 3.1 Pro plunged 10.8 points, Qwen3 Max plunged 11.8 points, Claude Sonnet 4.6 plunged 12.9 points, GPT-o3 plunged 13.7 points, and GPT-5.5 plunged 16.7 points in material constraint. The Smoke test provides small-sample single-day signals; these anomalies require more run data to verify their stability.
Key Changes
- GPT-o3: Leaderboard down 13.7 points, code execution -8.5 points, material constraint -20 points
- Claude Sonnet 4.6: Leaderboard down 12.9 points, code execution -8.5 points, material constraint -18.3 points
- Qwen3 Max: Leaderboard down 11.8 points, material constraint -27.2 points
- Gemini 3.1 Pro: Leaderboard down 10.8 points, code execution -7.8 points, material constraint -14.4 points
- Claude Opus 4.7: Leaderboard up 9.8 points, code execution +16.5 points
Signals to Watch
- Grok 4: Leaderboard plunged -8.8 points
- Gemini 3.1 Pro: Leaderboard plunged -10.8 points
- Qwen3 Max: Leaderboard plunged -11.8 points
- Claude Sonnet 4.6: Leaderboard plunged -12.9 points
- GPT-o3: Leaderboard plunged -13.7 points
- GPT-5.5: Material constraint plunged -16.7 points
When reading this type of Smoke briefing, the focus should be on two questions: first, whether a model has exposed the same type of weakness on multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large single-day swings in execution or constraint scores may come from question sampling or could be early signals of genuine regression, requiring follow-up runs for verification.
Data source: YZ Index | Run #300 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接