On 2026-08-17, the YZ Index Smoke quick test covered 11 models, with Claude Opus 4.7 and Grok 4 tied for the top score of the day at 96.99. Smoke is a daily 10-question quick test, suitable for observing short-term signals and not equivalent to Full weekly ranking conclusions.
This Smoke evaluation only covers the two main leaderboard dimensions: code execution and material constraints. The leaderboard formula is 0.55 × code execution + 0.45 × material constraints. Because the daily sample size is small, single-day scores are better suited as monitoring signals than as long-term conclusions about model capability.
Daily Rankings
| Rank | Model | Leaderboard | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | Claude Opus 4.7 | 96.99 | 100 | 93.3 | pass |
| #2 | Grok 4 | 96.99 | 100 | 93.3 | warn |
| #3 | Qwen3 Max | 95.34 | 97 | 93.3 | pass |
| #4 | Doubao Pro | 94.74 | 100 | 88.3 | pass |
| #5 | GPT-o3 | 94.02 | 100 | 86.7 | pass |
| #6 | Claude Sonnet 4.6 | 83.24 | 75 | 93.3 | pass |
| #7 | Gemini 3.1 Pro | 83.24 | 75 | 93.3 | pass |
| #8 | GPT-5.5 | 83.24 | 75 | 93.3 | pass |
| #9 | Gemini 2.5 Pro | 55.74 | 25 | 93.3 | pass |
| #10 | GLM-4.6 | 55.74 | 25 | 93.3 | pass |
| #11 | DeepSeek V4 Pro | 51.24 | 25 | 83.3 | pass |
Data Interpretation
Among today's top five models on the leaderboard, Claude Opus 4.7 and Grok 4 both tied at 96.99 with code execution 100 and material constraints 93.3, while Qwen3 Max scored 95.34 with code execution 97 and material constraints 93.3, indicating that a combination of perfect or near-perfect code execution with material constraints of 93.3 forms the top tier in the current sample. Doubao Pro reached 94.74 with code execution 100 and material constraints 88.3, and GPT-o3 reached 94.02 with code execution 100 and material constraints 86.7, showing that these models emphasize different balances between code execution and material constraints.
Compared with the previous run on the same basis, Doubao Pro's leaderboard score rose by 24.3 points, code execution by 33.3 points, and material constraints by 13.3 points; Qwen3 Max's leaderboard score rose by 20.3 points, code execution by 22 points, and material constraints by 18.3 points; Gemini 3.1 Pro's leaderboard score rose by 12.8 points, code execution by 8.3 points, and material constraints by 18.3 points; Claude Opus 4.7's leaderboard score rose by 10.7 points, code execution by 25 points, and material constraints fell by 6.7 points; DeepSeek V4 Pro's leaderboard score fell by 19.2 points, code execution by 41.7 points, and material constraints rose by 8.3 points. DeepSeek V4 Pro's sharp 19.2-point drop may stem from question-sampling fluctuation or may reflect real degradation, requiring confirmation in subsequent runs.
Mid- and low-tier models such as Gemini 2.5 Pro and GLM-4.6, with code execution 25 and material constraints 93.3, share the same leaderboard score of 55.74, while DeepSeek V4 Pro, with code execution 25 and material constraints 83.3, has a leaderboard score of 51.24, showing that when code execution is low, a high material-constraints score can still sustain a certain leaderboard position, but the gap from the top tier is significant. Smoke provides a small-sample single-day signal; the above observations are for that day's reference only.
Key Changes
- Doubao Pro: leaderboard +24.3, code execution +33.3, material constraints +13.3
- Qwen3 Max: leaderboard +20.3, code execution +22, material constraints +18.3
- DeepSeek V4 Pro: leaderboard -19.2, code execution -41.7, material constraints +8.3
- Gemini 3.1 Pro: leaderboard +12.8, code execution +8.3, material constraints +18.3
- Claude Opus 4.7: leaderboard +10.7, code execution +25, material constraints -6.7
Signals to Watch
- DeepSeek V4 Pro: leaderboard plunged -19.2 points
When reading Smoke briefs of this kind, the focus should be on two questions: first, whether a given model has exposed the same type of weakness for several consecutive days; second, whether its integrity rating has moved from pass to warn or fail. Large daily swings in execution or constraint scores may come from question sampling or may be early signals of real degradation, and require verification in subsequent runs.
Data source: YZ Index | Run #281 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接