On 2026-07-26, the YZ Index Smoke quick test covered 10 models, with DeepSeek V4 Pro ranking first with a score of 83.23. Smoke is a daily 10-question quick test suitable for observing short-term signals and is not equivalent to the Full weekly ranking conclusion.
This Smoke evaluation only covers the two main rank dimensions: Code Execution and Material Constraints. The main rank formula is 0.55 × Code Execution + 0.45 × Material Constraints. Due to the small daily sample size, single-day scores are more suitable as monitoring signals rather than making long-term conclusions about model capabilities.
Daily Ranking
| Rank | Model | Main Rank | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | DeepSeek V4 Pro | 83.23 | 69.5 | 100 | pass |
| #2 | Gemini 3.1 Pro | 81.11 | 86.1 | 75 | pass |
| #3 | Doubao Pro | 77.47 | 69.5 | 87.2 | pass |
| #4 | Gemini 2.5 Pro | 73.51 | 69.5 | 78.4 | pass |
| #5 | GPT-o3 | 69.48 | 44.5 | 100 | pass |
| #6 | Grok 4 | 69.48 | 44.5 | 100 | pass |
| #7 | Claude Sonnet 4.6 | 64.75 | 44.5 | 89.5 | pass |
| #8 | GPT-5.5 | 59.76 | 44.5 | 78.4 | pass |
| #9 | Claude Opus 4.7 | 55.73 | 19.5 | 100 | pass |
| #10 | Qwen3 Max | 43.13 | 19.5 | 72 | pass |
Data Interpretation
In today's YZ Index Smoke quick test, DeepSeek V4 Pro ranked first with a main rank of 83.23, with a structure of Code Execution 69.5 combined with Material Constraints 100, showing outstanding performance on the Material Constraints side. Gemini 3.1 Pro followed closely with a main rank of 81.11, Code Execution 86.1, and Material Constraints 75, reflecting a relative advantage on the Code Execution side. Doubao Pro scored a main rank of 77.47, with Code Execution 69.5 and Material Constraints 87.2, presenting a relatively balanced combination of the two indicators. The top-ranking models each focused on different strengths between Code Execution and Material Constraints, forming the current main rank pattern.
Compared to the previous same-criteria run, Claude Sonnet 4.6's main rank dropped by 32.2 points, Code Execution dropped by 50 points, and Material Constraints dropped by 10.5 points. Grok 4's main rank dropped by 27.5 points, Code Execution dropped by 50 points. GPT-5.5's main rank dropped by 27.4 points, Code Execution dropped by 50 points. Claude Opus 4.7's main rank dropped by 22.7 points, Code Execution dropped by 50 points, while Material Constraints increased by 10.7 points. GPT-o3's main rank dropped by 11.4 points, Code Execution dropped by 50 points, while Material Constraints increased by 35.7 points. These changes are mainly concentrated in the Code Execution indicator, which may be due to question sampling fluctuations or real performance differences of the model under specific constraints. Subsequent runs are needed to review and confirm signal stability.
As Smoke is a small-sample single-day signal, the above analysis of score structure only reflects the characteristics of the day's data and does not make long-term judgments about the overall capabilities of the models.
Key Changes
- Claude Sonnet 4.6: Main rank down 32.2 points, Code Execution -50 points, Material Constraints -10.5 points
- Grok 4: Main rank down 27.5 points, Code Execution -50 points
- GPT-5.5: Main rank down 27.4 points, Code Execution -50 points
- Claude Opus 4.7: Main rank down 22.7 points, Code Execution -50 points, Material Constraints +10.7 points
- GPT-o3: Main rank down 11.4 points, Code Execution -50 points, Material Constraints +35.7 points
Signals to Watch
- No reportable anomaly signals were retained this time.
When reading such Smoke briefs, the focus should be on two issues: first, whether a model has been exposing the same type of weakness for consecutive days; second, whether the Integrity rating has moved from pass to warn or fail. Large daily changes in execution or constraint scores may come from question sampling or be early signals of real degradation, requiring subsequent runs to review.
Data Source: YZ Index | Run #246 | View Raw Data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接