On 2026-07-23, the YZ Index Smoke Quick Test covered 11 models, with Claude Opus 4.7 ranking first at 96.99. Smoke is a daily 10-question quick test suitable for observing short-term signals, but does not equate to the conclusions of the Full weekly ranking.
This Smoke test only covers two main dimensions of the primary benchmark: Code Execution and Material Constraint. The primary benchmark formula is 0.55 × Code Execution + 0.45 × Material Constraint. Due to the small daily sample size, single-day scores are best used as monitoring signals rather than long-term conclusions about model capabilities.
Daily Ranking
| Rank | Model | Primary Benchmark | Code Execution | Material Constraint | Integrity |
|---|---|---|---|---|---|
| #1 | Claude Opus 4.7 | 96.99 | 100 | 93.3 | pass |
| #2 | Doubao Pro | 88.8 | 100 | 75.1 | pass |
| #3 | Qwen3 Max | 82.03 | 83.6 | 80.1 | pass |
| #4 | Grok 4 | 80.28 | 84.6 | 75 | pass |
| #5 | GPT-o3 | 75 | 75 | 75 | pass |
| #6 | GPT-5.5 | 72.62 | 71.9 | 73.5 | pass |
| #7 | Claude Sonnet 4.6 | 72.03 | 75 | 68.4 | pass |
| #8 | DeepSeek V4 Pro | 72.03 | 75 | 68.4 | pass |
| #9 | Gemini 2.5 Pro | 69.49 | 50 | 93.3 | pass |
| #10 | Gemini 3.1 Pro | 69.49 | 50 | 93.3 | pass |
| #11 | GLM-4.6 | 55.74 | 25 | 93.3 | fail |
Data Interpretation
In today's YZ Index Smoke Quick Test, Claude Opus 4.7 leads with a primary benchmark of 96.99. Its Code Execution score of 100 and Material Constraint score of 93.3 indicate strong performance in both dimensions. Doubao Pro scored 88.8 on the primary benchmark, also achieving 100 on Code Execution but 75.1 on Material Constraint, reflecting outstanding performance in code execution. Qwen3 Max scored 82.03 on the primary benchmark, with Code Execution at 83.6 and Material Constraint at 80.1, forming a relatively balanced structure. Lower-ranked Gemini 2.5 Pro and Gemini 3.1 Pro both scored 69.49 on the primary benchmark, with Code Execution at 50 and Material Constraint at 93.3, showing a clear advantage in material constraint. GLM-4.6 scored 55.74 on the primary benchmark, with Code Execution at 25, Material Constraint at 93.3, and an Integrity rating of fail.
Compared to the previous run with the same scope, Claude Sonnet 4.6 dropped 24.7 points on the primary benchmark, with Code Execution down 19 points and Material Constraint down 31.6 points. DeepSeek V4 Pro also dropped 24.7 points on the primary benchmark, with Code Execution down 19 points and Material Constraint down 31.6 points. GPT-5.5 dropped 24.1 points on the primary benchmark, with Code Execution down 22.1 points and Material Constraint down 26.5 points. GLM-4.6 dropped 18.3 points on the primary benchmark, with Code Execution down 72 points and Material Constraint up 18.3 points. Grok 4 dropped 18.1 points on the primary benchmark, with Code Execution down 12.4 points and Material Constraint down 25 points. These score changes may be due to question sampling fluctuations or actual model degradation, requiring verification in subsequent runs.
The Smoke Quick Test provides small-sample single-day signals. The above observations only reflect the data structure characteristics of the day and should not be used for long-term conclusions.
Key Changes
- Claude Sonnet 4.6: Primary benchmark down 24.7 pts, Code Execution -19 pts, Material Constraint -31.6 pts, Integrity warn→pass
- DeepSeek V4 Pro: Primary benchmark down 24.7 pts, Code Execution -19 pts, Material Constraint -31.6 pts
- GPT-5.5: Primary benchmark down 24.1 pts, Code Execution -22.1 pts, Material Constraint -26.5 pts
- GLM-4.6: Primary benchmark down 18.3 pts, Code Execution -72 pts, Material Constraint +18.3 pts
- Grok 4: Primary benchmark down 18.1 pts, Code Execution -12.4 pts, Material Constraint -25 pts
Signals to Watch
- GLM-4.6: Integrity rating is fail today (based on today's Smoke data).
When reading these Smoke briefs, the focus should be on two questions: first, whether a model has been repeatedly exposed to the same weakness over consecutive days; second, whether the Integrity rating has moved from pass to warn or fail. Large daily fluctuations in execution or constraint scores may stem from question sampling or be early signals of actual degradation, and require verification in subsequent runs.
Data source: YZ Index | Run #243 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接