2026-07-31 YZ Index Smoke quick test covered 10 models, with DeepSeek V4 Pro scoring 96.94 to top the daily rankings. Smoke is a daily 10-question quick test suitable for observing short-term signals and is not equivalent to the Full weekly ranking conclusions.
This Smoke evaluation only covers two main ranking dimensions: code execution and material constraints. The main ranking formula is 0.55 × code execution + 0.45 × material constraints. Due to the small daily sample size, the single-day score is more suitable as a monitoring signal rather than a long-term conclusion about model capabilities.
Daily Rankings
| Rank | Model | Overall | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | DeepSeek V4 Pro | 96.94 | 100 | 93.2 | pass |
| #2 | GPT-o3 | 93.16 | 95 | 90.9 | pass |
| #3 | Claude Opus 4.7 | 92.85 | 100 | 84.1 | pass |
| #4 | GPT-5.5 | 86.63 | 95 | 76.4 | pass |
| #5 | Grok 4 | 77.72 | 72.5 | 84.1 | pass |
| #6 | Gemini 3.1 Pro | 76.35 | 70 | 84.1 | pass |
| #7 | Doubao Pro | 75 | 75 | 75 | pass |
| #8 | Claude Sonnet 4.6 | 74.04 | 100 | 42.3 | pass |
| #9 | Qwen3 Max | 72.34 | 92.5 | 47.7 | pass |
| #10 | Gemini 2.5 Pro | 69.26 | 72 | 65.9 | pass |
Data Interpretation
In today's YZ Index Smoke quick test, DeepSeek V4 Pro achieved an overall score of 96.94 with a combination of 100 in code execution and 93.2 in material constraints, while GPT-o3 scored 93.16 with 95 in code execution and 90.9 in material constraints. Claude Opus 4.7 also scored 100 in code execution but 84.1 in material constraints, resulting in an overall score of 92.85. Top models generally show a balanced or strongly leaning combination of code execution and material constraints, whereas Claude Sonnet 4.6 (100 code execution but 42.3 material constraints) and Qwen3 Max (92.5 code execution but 47.7 material constraints) exhibit significant disparity.
Compared to the previous run with the same criteria, DeepSeek V4 Pro saw a +20.1 point overall increase, +25 points in code execution, and +14.1 points in material constraints; GPT-o3 increased by +13.3 points overall and +20 points in code execution; Qwen3 Max increased by +11.8 points overall but dropped -20 points in material constraints; Gemini 2.5 Pro increased by +20.9 points in material constraints. These fluctuations may be due to question sampling variance or could indicate real degradation, requiring confirmation in subsequent runs.
Grok 4's code execution plummeted by -19.5 points, and Qwen3 Max's material constraints dropped sharply by -20 points, forming anomalous signals. GLM-4.6 had incomplete data due to API failure and has been automatically rescheduled for a rerun; it is not included in this ranking. As a small-sample single-day signal, the above changes are for reference only and do not constitute long-term conclusions.
Key Changes
- DeepSeek V4 Pro: Overall +20.1, Code Execution +25, Material Constraints +14.1
- GPT-o3: Overall +13.3, Code Execution +20, Material Constraints +5
- Qwen3 Max: Overall +11.8, Code Execution +37.8, Material Constraints -20
- Gemini 2.5 Pro: Overall +8.5, Material Constraints +20.9
- Claude Opus 4.7: Overall +6.3, Material Constraints +14.1
Signals to Watch
- Grok 4: Code Execution plummeted by -19.5 points
- Qwen3 Max: Material Constraints plummeted by -20 points
- GLM-4.6: Incomplete data (missing execution, grounding, judgment, integrity, communication dimensions due to API failure/timeout); automatically rescheduled for rerun, not included in this ranking
When reading such Smoke briefs, the focus should be on two questions: first, whether a model consistently exposes the same weakness over consecutive days; second, whether the integrity rating has shifted from pass to warn or fail. Large day-to-day fluctuations in execution or constraints scores may be due to question sampling or early signs of real degradation, requiring subsequent runs for verification.
Data source: YZ Index (YZ Index) | Run #255 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接