On 2026-08-11, the YZ Index Smoke quick test covered 10 models, with Gemini 3.1 Pro ranking first that day at 94.74 points. Smoke is a daily 10-question quick test suited for observing short-term signals; it is not equivalent to the Full weekly ranking conclusion.
This Smoke evaluation only covers two main ranking dimensions: code execution and material constraints. The main ranking formula is 0.55 × code execution + 0.45 × material constraints. Given the small daily sample size, single-day scores are better used as monitoring signals rather than as long-term verdicts on model capabilities.
Daily Ranking
| Rank | Model | Main Score | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | Gemini 3.1 Pro | 94.74 | 100 | 88.3 | pass |
| #2 | Claude Opus 4.7 | 92.94 | 100 | 84.3 | pass |
| #3 | DeepSeek V4 Pro | 87.99 | 100 | 73.3 | pass |
| #4 | Grok 4 | 87.93 | 95.8 | 78.3 | pass |
| #5 | Gemini 2.5 Pro | 86.28 | 92.8 | 78.3 | pass |
| #6 | Claude Sonnet 4.6 | 85.74 | 100 | 68.3 | pass |
| #7 | GPT-o3 | 84.41 | 94.8 | 71.7 | pass |
| #8 | GPT-5.5 | 83.17 | 100 | 62.6 | pass |
| #9 | Qwen3 Max | 81.56 | 87.5 | 74.3 | pass |
| #10 | GLM-4.6 | 74 | 75 | 78.3 | fail |
Key Changes
- Claude Sonnet 4.6: Main Score up 30.5 points, Code Execution +50 points, Material Constraints +6.6 points
- Gemini 3.1 Pro: Main Score up 28 points, Code Execution +33.3 points, Material Constraints +21.6 points
- GPT-5.5: Main Score up 24.9 points, Code Execution +25 points, Material Constraints +24.7 points
- Claude Opus 4.7: Main Score up 20.6 points, Code Execution +25 points, Material Constraints +15.3 points
- Qwen3 Max: Main Score up 16.6 points, Code Execution +12.5 points, Material Constraints +21.5 points
Signals to Watch
- GLM-4.6: Integrity rating downgraded to Fail (pass→fail)
- Doubao Pro: Incomplete data (missing execution, material-constraint, judgment, integrity, and communication dimensions; API failure/timeout). Automatic re-run initiated; not ranked in this round.
When reading Smoke briefings of this kind, the focus should be on two questions: first, whether a model has exposed the same type of weakness on multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large single-day swings in execution or constraint scores may stem from question sampling, or they may be early signals of genuine regression, requiring review in subsequent runs.
Data source: YZ Index | Run #274 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接