The 2026-09-29 YZ Index Smoke quick test covered 10 models, with Claude Opus 4.7 and GPT-5.5 tied for first place that day at 80.2. Smoke is a daily 10-question quick test, suitable for observing short-term signals, and is not equivalent to Full weekly leaderboard conclusions.
This Smoke evaluation covered only two Main Leaderboard dimensions: code execution and material constraints. The Main Leaderboard formula is 0.55 × code execution + 0.45 × material constraints. Because the daily sample size is small, single-day scores are better treated as monitoring signals rather than as long-term verdicts on model capability.
Rankings for the Day
| Rank | Model | Main Leaderboard | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | Claude Opus 4.7 | 80.2 | 100 | 56 | pass |
| #2 | GPT-5.5 | 80.2 | 100 | 56 | pass |
| #3 | Doubao Pro | 76.25 | 96.9 | 51 | warn |
| #4 | Qwen3 Max | 75.91 | 92.2 | 56 | warn |
| #5 | GPT-o3 | 73.78 | 93.8 | 49.3 | warn |
| #6 | Grok 4 | 69.98 | 94.5 | 40 | pass |
| #7 | Claude Sonnet 4.6 | 66.3 | 87 | 41 | pass |
| #8 | Gemini 3.1 Pro | 54.95 | 50 | 61 | pass |
| #9 | DeepSeek V4 Pro | 45.7 | 25 | 71 | warn |
| #10 | Gemini 2.5 Pro | 41.2 | 25 | 61 | pass |
Data Interpretation
Today's top two on the Main Leaderboard, Claude Opus 4.7 and GPT-5.5, each scored 80.2 overall, with code execution at 100 and material constraints at 56 for both, indicating that the two form a balanced high-end combination of code execution and material constraints. Doubao Pro scored 76.25 on the Main Leaderboard, with code execution at 96.9 and material constraints at 51, while Qwen3 Max scored 75.91 overall, pairing code execution of 92.2 with material constraints of 56. Overall, the leading models display a structural pattern of relatively strong code execution and moderate material constraints. Gemini 3.1 Pro scored 54.95 on the Main Leaderboard, with code execution at 50 and material constraints at 61, while DeepSeek V4 Pro scored 45.7 overall, with code execution at 25 and material constraints at 71, reflecting a different combination in which material constraints are relatively stronger.
Claude Sonnet 4.6 fell 25.5 points on the Main Leaderboard compared with the previous run on the same basis, with code execution down 13 points and material constraints down 40.8 points; DeepSeek V4 Pro fell 25.2 points overall, with code execution down 50 points and material constraints up 5.1 points; Gemini 2.5 Pro fell 22.9 points overall, with code execution down 45 points; Grok 4 fell 21.8 points overall, with code execution down 5.5 points and material constraints down 41.8 points; and Qwen3 Max fell 15.9 points overall, with material constraints down 34.9 points. These score changes may stem from variation in question sampling, or they may reflect real differences in single-day performance, and they need to be confirmed by follow-up runs. As a small-sample, single-day signal, Smoke's interpretation above applies only to today's data structure and does not constitute a long-term judgment.
Key Changes
- Claude Sonnet 4.6: Main Leaderboard down 25.5 points, code execution −13 points, material constraints −40.8 points
- DeepSeek V4 Pro: Main Leaderboard down 25.2 points, code execution −50 points, material constraints +5.1 points
- Gemini 2.5 Pro: Main Leaderboard down 22.9 points, code execution −45 points
- Grok 4: Main Leaderboard down 21.8 points, code execution −5.5 points, material constraints −41.8 points
- Qwen3 Max: Main Leaderboard down 15.9 points, material constraints −34.9 points
Signals to Watch
- No anomaly signals eligible for publication were retained in this test.
When reading this kind of Smoke brief, the focus should be on two questions: first, whether a given model exposes the same type of weakness on several consecutive days; second, whether its integrity rating moves from pass into warn or fail. Large single-day swings in execution or constraint scores may come from question sampling, or they may be early signals of genuine degradation, and they need to be confirmed by follow-up runs.
Data source: YZ Index | Run #344 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接