On 2026-07-29, the Winzheng YZ Index Smoke quick test covered 10 models, with Grok 4 leading the day at 89.3 points. Smoke is a daily 10-question quick test designed for observing short-term signals and is not equivalent to the full-week benchmark conclusions.
This Smoke evaluation only covers two main dimensions: Code Execution and Material Constraint. The main formula is 0.55 × Code Execution + 0.45 × Material Constraint. Due to the small daily sample size, single-day scores are better used as monitoring signals rather than long-term conclusions about model capabilities.
Daily Rankings
| Rank | Model | Main Score | Code Execution | Material Constraint | Integrity |
|---|---|---|---|---|---|
| #1 | Grok 4 | 89.3 | 97.8 | 78.9 | pass |
| #2 | DeepSeek V4 Pro | 83.53 | 100 | 63.4 | warn |
| #3 | Doubao Pro | 81.88 | 97 | 63.4 | pass |
| #4 | GPT-o3 | 79.43 | 72.5 | 87.9 | pass |
| #5 | Gemini 3.1 Pro | 77.7 | 78.6 | 76.6 | pass |
| #6 | Claude Opus 4.7 | 76.76 | 75 | 78.9 | pass |
| #7 | Claude Sonnet 4.6 | 70.77 | 75 | 65.6 | pass |
| #8 | Qwen3 Max | 68.05 | 72.5 | 62.6 | pass |
| #9 | GPT-5.5 | 66.68 | 75 | 56.5 | pass |
| #10 | Gemini 2.5 Pro | 58.91 | 50 | 69.8 | pass |
Data Interpretation
In today's Winzheng YZ Index Smoke quick test, top models showed clear differentiation in the combination of Code Execution and Material Constraint. Grok 4 leads with a main score of 89.3, leveraging a balanced advantage of 97.8 in Code Execution and 78.9 in Material Constraint; DeepSeek V4 Pro achieves a perfect 100 in Code Execution but only 63.4 in Material Constraint, placing second with a main score of 83.53; Doubao Pro follows closely with Code Execution at 97 and Material Constraint at 63.4, yielding a main score of 81.88. In contrast, GPT-o3 scores 72.5 in Code Execution but achieves a main score of 79.43 due to its strong Material Constraint of 87.9; Gemini 3.1 Pro records 78.6 in Code Execution and 76.6 in Material Constraint, with a main score of 77.7, indicating that models with stronger Material Constraint remain competitive in the overall ranking.
Notable changes: Claude Opus 4.7 saw its main score rise by 20.3, Code Execution up by 50, and Material Constraint down by 16.1; GPT-5.5's main score dropped by 27, Code Execution down by 25, and Material Constraint down by 29.4; Gemini 3.1 Pro's main score fell by 22.3, Code Execution down by 21.4, and Material Constraint down by 23.4; Gemini 2.5 Pro's main score decreased by 18.7, Code Execution down by 20.8, and Material Constraint down by 16.1; GPT-o3's main score declined by 18.3, Code Execution down by 27.5, and Material Constraint down by 7.1. These single-day fluctuations may stem from question sampling variation or reflect short-term changes in model performance, requiring subsequent runs with consistent settings for verification.
As a small-sample single-day signal, this data provides only real-time reference. No abnormal signals were recorded. Interpretation remains restrained to avoid over-extrapolation.
Key Changes
- GPT-5.5: Main score down by 27 points, Code Execution -25 points, Material Constraint -29.4 points
- Gemini 3.1 Pro: Main score down by 22.3 points, Code Execution -21.4 points, Material Constraint -23.4 points
- Claude Opus 4.7: Main score up by 20.3 points, Code Execution +50 points, Material Constraint -16.1 points
- Gemini 2.5 Pro: Main score down by 18.7 points, Code Execution -20.8 points, Material Constraint -16.1 points
- GPT-o3: Main score down by 18.3 points, Code Execution -27.5 points, Material Constraint -7.1 points
Signals to Watch
- No publishable abnormal signals were retained this time.
When reading such Smoke briefs, focus on two questions: first, whether a model has repeatedly exposed the same weakness over consecutive days; second, whether the integrity rating has shifted from pass to warn or fail. Large daily fluctuations in execution or constraint scores may stem from question sampling or could be early signs of genuine degradation, requiring verification in subsequent runs.
Data Source: Winzheng YZ Index | Run #252 | View Raw Data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接