On 2026-07-27, the YZ Index Smoke quick test covered 11 models, with GPT-o3 ranking first on the day with a score of 91.29. Smoke is a daily quick test of 10 questions, suitable for observing short-term signals, and is not equivalent to the Full weekly ranking conclusions.
This Smoke evaluation only covers the two main ranking dimensions: Code Execution and Material Constraints. The main ranking formula is 0.55 × Code Execution + 0.45 × Material Constraints. Due to the small daily sample size, single-day scores are more suitable as monitoring signals rather than long-term conclusions about model capabilities.
Daily Rankings
| Rank | Model | Main Ranking | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | GPT-o3 | 91.29 | 97 | 84.3 | pass |
| #2 | Gemini 3.1 Pro | 89.04 | 97 | 79.3 | pass |
| #3 | GPT-5.5 | 86.29 | 97 | 73.2 | pass |
| #4 | DeepSeek V4 Pro | 85.69 | 100 | 68.2 | pass |
| #5 | Doubao Pro | 84.94 | 97 | 70.2 | pass |
| #6 | Gemini 2.5 Pro | 84.28 | 95.8 | 70.2 | pass |
| #7 | Claude Sonnet 4.6 | 80.44 | 97 | 60.2 | pass |
| #8 | GLM-4.6 | 75.69 | 72 | 80.2 | warn |
| #9 | Grok 4 | 75.15 | 79.2 | 70.2 | pass |
| #10 | Claude Opus 4.7 | 66.04 | 47 | 89.3 | pass |
| #11 | Qwen3 Max | 62.07 | 59.5 | 65.2 | pass |
Data Interpretation
Today’s YZ Index Smoke quick test shows that top models have different emphases in the combination of Code Execution and Material Constraints. GPT-o3 has a main ranking of 91.29, with Code Execution 97 and Material Constraints 84.3, showing a relatively balanced capability; Gemini 3.1 Pro has a main ranking of 89.04, Code Execution 97, Material Constraints 79.3, also relying on high Code Execution scores; DeepSeek V4 Pro has Code Execution 100, Material Constraints 68.2, main ranking 85.69, reflecting a structure with clear Code Execution advantage but weaker Material Constraints; GPT-5.5 and Doubao Pro both have Code Execution 97, Material Constraints 73.2 and 70.2 respectively, overall ranking high on the main ranking.
Among significant changes, GPT-5.5 main ranking +26.5 points, Code Execution +52.5 points; GPT-o3 main ranking +21.8 points, Code Execution +52.5 points; Qwen3 Max main ranking +18.9 points, Code Execution +40 points; Claude Sonnet 4.6 main ranking +15.7 points, Code Execution +52.5 points. These improvements were accompanied by varying degrees of decline in Material Constraints. Abnormal signals include GPT-o3 Material Constraints plummeting -15.7 points, DeepSeek V4 Pro Material Constraints plummeting -31.8 points, Doubao Pro Material Constraints plummeting -17 points, Claude Sonnet 4.6 Material Constraints plummeting -29.3 points, Grok 4 Material Constraints plummeting -29.8 points, and GLM-4.6 Integrity rating downgraded to warn, which may be due to question sampling fluctuations or genuine degradation, requiring confirmation in subsequent runs. Smoke is a small-sample single-day signal, and interpretations should remain restrained.
Major Changes
- GPT-5.5: Main ranking +26.5 points, Code Execution +52.5 points, Material Constraints -5.2 points
- GPT-o3: Main ranking +21.8 points, Code Execution +52.5 points, Material Constraints -15.7 points
- Qwen3 Max: Main ranking +18.9 points, Code Execution +40 points, Material Constraints -6.8 points
- GLM-4.6: Main ranking +18.5 points, Code Execution +27.5 points, Material Constraints +7.4 points, Integrity fail→warn
- Claude Sonnet 4.6: Main ranking +15.7 points, Code Execution +52.5 points, Material Constraints -29.3 points
Signals to Watch
- GPT-o3: Material Constraints plummeted -15.7 points
- DeepSeek V4 Pro: Material Constraints plummeted -31.8 points
- Doubao Pro: Material Constraints plummeted -17 points
- Claude Sonnet 4.6: Material Constraints plummeted -29.3 points
- GLM-4.6: Integrity rating downgraded to Fail (fail→warn)
- Grok 4: Material Constraints plummeted -29.8 points
When reading this type of Smoke brief, the focus should be on two questions: first, whether a certain model has exposed the same type of weakness for consecutive days; second, whether the Integrity rating has moved from pass to warn or fail. Large daily changes in Execution or Constraint scores may come from question sampling or be early signals of genuine degradation, requiring confirmation in subsequent runs.
Data source: YZ Index | Run #248 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接