On 2026-07-20, the YZ Index Smoke Quick Test covered 11 models, with Claude Opus 4.7 scoring 100 points to top the daily rankings. Smoke is a daily 10-question quick test, suitable for observing short-term signals, and is not equivalent to the Full weekly ranking conclusions.
This Smoke evaluation only covers two main ranking dimensions: code execution and material constraints. The main ranking formula is 0.55 × Code Execution + 0.45 × Material Constraints. Due to the small daily sample size, single-day scores are better suited as monitoring signals rather than long-term conclusions about model capabilities.
Daily Rankings
| Rank | Model | Main Score | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | Claude Opus 4.7 | 100 | 100 | 100 | pass |
| #2 | Doubao Pro | 83.73 | 75 | 94.4 | pass |
| #3 | GPT-o3 | 82.7 | 75 | 92.1 | pass |
| #4 | Gemini 3.1 Pro | 81.93 | 75 | 90.4 | pass |
| #5 | Grok 4 | 80.81 | 75 | 87.9 | warn |
| #6 | DeepSeek V4 Pro | 80.27 | 75 | 86.7 | warn |
| #7 | GPT-5.5 | 79.91 | 75 | 85.9 | pass |
| #8 | Gemini 2.5 Pro | 69.98 | 50 | 94.4 | pass |
| #9 | Qwen3 Max | 67.31 | 65.6 | 69.4 | warn |
| #10 | Claude Sonnet 4.6 | 64.39 | 71.9 | 55.2 | pass |
| #11 | GLM-4.6 | 58.75 | 25 | 100 | warn |
Data Interpretation
Among today's top four models in the main ranking, Claude Opus 4.7 achieved a main score of 100 with both code execution and material constraints at 100, showing a balanced high level. Doubao Pro's main score of 83.73 came from code execution at 75 and material constraints at 94.4; GPT-o3's main score of 82.7 corresponded to code execution at 75 and material constraints at 92.1; Gemini 3.1 Pro's main score of 81.93 came from code execution at 75 and material constraints at 90.4. This structure indicates that combinations with relatively stronger material constraints contributed higher overall scores in the current sample.
Doubao Pro's main score increased by 34.5 points compared to the previous run under the same conditions, mainly driven by a 50-point rise in code execution. GLM-4.6's main score rose 27.3 points, accompanied by a 60.7-point increase in material constraints. Gemini 3.1 Pro's main score increased by 25.5 points, with code execution up 25 points and material constraints up 26.1 points. Regarding anomalies, Gemini 2.5 Pro's code execution saw a sharp drop of -24.6 points, Qwen3 Max's main score plunged -14.9 points, and Claude Sonnet 4.6's main score plummeted -25.6 points. These single-day changes may stem from question sampling fluctuations or could indicate real performance degradation, requiring confirmation from subsequent runs under the same conditions.
The Smoke test is a small-sample single-day signal; the above observations only reflect the day's data distribution and do not make judgments on long-term model stability.
Key Changes
- Doubao Pro: Main score +34.5, Code Execution +50, Material Constraints +15.6
- GLM-4.6: Main score +27.3, Material Constraints +60.7
- Claude Sonnet 4.6: Main score -25.6, Code Execution -27.3, Material Constraints -23.6
- Gemini 3.1 Pro: Main score +25.5, Code Execution +25, Material Constraints +26.1
- DeepSeek V4 Pro: Main score +17.3, Code Execution +25, Material Constraints +7.9, Integrity pass→warn
Signals to Watch
- Gemini 2.5 Pro: Code Execution plunged -24.6 points
- Qwen3 Max: Main score plunged -14.9 points
- Claude Sonnet 4.6: Main score plunged -25.6 points
When reading such Smoke briefs, the focus should be on two questions: first, whether a particular model exposes the same type of weakness on consecutive days; second, whether the integrity rating has shifted from pass to warn or fail. Large daily fluctuations in execution or constraint scores may stem from question sampling or could be early signals of real degradation, requiring confirmation from subsequent runs.
Data source: YZ Index | Run #238 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接