On July 25, 2026, the YZ Index Smoke test covered 11 models, with Claude Sonnet 4.6 and Grok 4 both scoring 96.98, tying for first place. Smoke is a daily 10-question quick test designed to observe short-term signals and is not equivalent to the full weekly ranking conclusions.
This Smoke test only covers the two main ranking dimensions of code execution and material constraints. The main ranking formula is 0.55 × Code Execution + 0.45 × Material Constraints. Due to the small daily sample size, single-day scores are more suitable as monitoring signals rather than long-term conclusions about model capabilities.
Daily Rankings
| Rank | Model | Main Score | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | Claude Sonnet 4.6 | 96.98 | 94.5 | 100 | pass |
| #2 | Grok 4 | 96.98 | 94.5 | 100 | pass |
| #3 | DeepSeek V4 Pro | 92.16 | 94.5 | 89.3 | pass |
| #4 | GPT-5.5 | 87.17 | 94.5 | 78.2 | pass |
| #5 | Gemini 3.1 Pro | 80.91 | 94.5 | 64.3 | pass |
| #6 | GPT-o3 | 80.91 | 94.5 | 64.3 | pass |
| #7 | Gemini 2.5 Pro | 80.73 | 94.5 | 63.9 | pass |
| #8 | Claude Opus 4.7 | 78.41 | 69.5 | 89.3 | pass |
| #9 | Doubao Pro | 71.98 | 69.5 | 75 | pass |
| #10 | GLM-4.6 | 64.66 | 44.5 | 89.3 | pass |
| #11 | Qwen3 Max | 53.41 | 44.5 | 64.3 | pass |
Data Interpretation
Today's top two models, Claude Sonnet 4.6 and Grok 4, both scored 96.98, with code execution at 94.5 and material constraints at 100 for both, showing a balanced high-level combination in both capabilities. DeepSeek V4 Pro ranked third with 92.16, with the same code execution score of 94.5 but a slightly lower material constraints score of 89.3. GPT-5.5 scored 87.17 on the main ranking, with code execution at 94.5 and material constraints at 78.2, reflecting a structural advantage in code execution over material constraints.
Several models showed significant positive changes. Claude Sonnet 4.6 rose 38.3 points on the main ranking, 44.5 points in code execution, and 30.7 points in material constraints. DeepSeek V4 Pro rose 23.5 points on the main ranking, 19.5 points in code execution, and 28.4 points in material constraints. Gemini 3.1 Pro rose 21.4 points on the main ranking and 36.2 points in code execution. These increases come from a single-day small-sample Smoke test and may result from question sampling fluctuations or model performance variations under specific constraints, requiring subsequent runs under the same methodology to confirm stability.
Overall, top models tend to show complementary strengths between code execution and material constraints, while some mid-range models like GLM-4.6 (code execution 44.5, material constraints 89.3) and Qwen3 Max (code execution 44.5, material constraints 64.3) exhibit clear structural differences. All interpretations are based on today's data and do not constitute a judgment on long-term performance.
Key Changes
- Claude Sonnet 4.6: Main score +38.3, Code Execution +44.5, Material Constraints +30.7
- DeepSeek V4 Pro: Main score +23.5, Code Execution +19.5, Material Constraints +28.4
- Gemini 3.1 Pro: Main score +21.4, Code Execution +36.2
- GPT-o3: Main score +16, Code Execution +19.5, Material Constraints +11.7
- GPT-5.5: Main score +13.3, Code Execution +19.5, Material Constraints +5.6
Signals to Monitor
- No abnormal signals were retained for release this time.
When reading Smoke briefs like this, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether the integrity rating has shifted from pass to warn or fail. Large daily fluctuations in execution or constraint scores may result from question sampling or represent early signs of real degradation, requiring subsequent runs for verification.
Data source: YZ Index | Run #245 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接