On 2026-08-05, the YZ Index Smoke quick test covered 9 models, with DeepSeek V4 Pro, GPT-5.5, and GPT-o3 tying for first place at 80.52 points. Smoke is a daily 10-question quick test suitable for observing short-term signals and is not equivalent to the Full weekly ranking conclusion.
This Smoke evaluation only covers two main ranking dimensions: code execution and material constraints. The main ranking formula is 0.55 × Code Execution + 0.45 × Material Constraints. Due to the small daily sample size, single-day scores are better used as monitoring signals rather than as long-term conclusions about model capabilities.
Daily Ranking
| Rank | Model | Main Ranking | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | DeepSeek V4 Pro | 80.52 | 100 | 56.7 | pass |
| #2 | GPT-5.5 | 80.52 | 100 | 56.7 | pass |
| #3 | GPT-o3 | 80.52 | 100 | 56.7 | pass |
| #4 | Claude Sonnet 4.6 | 74.09 | 96.9 | 46.2 | pass |
| #5 | Gemini 2.5 Pro | 73.74 | 81.3 | 64.5 | pass |
| #6 | Claude Opus 4.7 | 71.94 | 75 | 68.2 | pass |
| #7 | Gemini 3.1 Pro | 55.52 | 75 | 31.7 | pass |
| #8 | Grok 4 | 55.52 | 75 | 31.7 | pass |
| #9 | Qwen3 Max | 49.72 | 71.9 | 22.6 | pass |
Data Interpretation
In today's YZ Index Smoke quick test, DeepSeek V4 Pro, GPT-5.5, and GPT-o3 tied on the main ranking at 80.52, each scoring 100 in code execution and 56.7 in material constraints, showing a structural combination of perfect code execution with moderate material constraints. Claude Sonnet 4.6 scored 74.09 on the main ranking, with 96.9 in code execution and 46.2 in material constraints, also exhibiting strong code execution and weaker material constraints. Gemini 2.5 Pro scored 73.74 on the main ranking, with 81.3 in code execution and 64.5 in material constraints, forming a relatively balanced but overall lower combination across both metrics. Claude Opus 4.7 scored 71.94 on the main ranking, with 75 in code execution and 68.2 in material constraints, where material constraints exceed code execution, forming another type of combination.
Compared with the previous run under the same methodology, Qwen3 Max fell 35.1 points on the main ranking, with code execution down 20.8 points and material constraints down 52.6 points; Grok 4 fell 28.8 points on the main ranking, with code execution down 25 points and material constraints down 33.5 points; Gemini 3.1 Pro fell 27.6 points on the main ranking, with code execution down 22 points and material constraints down 34.4 points. These declines are concentrated in material constraints and code execution, and subsequent runs are needed to distinguish between question-sampling fluctuations and genuine regression. The Smoke quick test is a small-sample single-day signal; the current data only reflects that day's performance and does not constitute a basis for long-term judgment.
Overall, top models largely rely on high code execution scores to support their main ranking, while material constraint scores are generally lower, creating a clear structural disparity. The synchronized decline among models showing anomalies suggests possible single-day test fluctuation, requiring multiple rounds of retesting to confirm stability.
Key Changes
- Qwen3 Max: Main ranking down 35.1 points; code execution -20.8; material constraints -52.6
- Grok 4: Main ranking down 28.8 points; code execution -25; material constraints -33.5
- Gemini 3.1 Pro: Main ranking down 27.6 points; code execution -22; material constraints -34.4
- Claude Opus 4.7: Main ranking down 17.5 points; code execution -22; material constraints -12
- Gemini 2.5 Pro: Main ranking down 15.8 points; code execution -18.7; material constraints -12.3
Signals to Watch
- No publishable anomaly signals were retained this time.
When reading this type of Smoke briefing, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large daily swings in execution or constraint scores may stem from question sampling or may be early signals of genuine regression, requiring subsequent run verification.
Data source: YZ Index | Run #262 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接