On 2026-08-09, the YZ Index Smoke quick test covered 9 models, with GPT-o3 ranking first on the day at 95.91 points. Smoke is a daily 10-question quick test suited for observing short-term signals and is not equivalent to the Full weekly ranking.
This Smoke evaluation only covers two main ranking dimensions: code execution and material constraint. The main ranking formula is 0.55 × code execution + 0.45 × material constraint. Given the small daily sample size, single-day scores are better used as monitoring signals rather than long-term conclusions about model capabilities.
Daily Ranking
| Rank | Model | Main Score | Code Execution | Material Constraint | Integrity |
|---|---|---|---|---|---|
| #1 | GPT-o3 | 95.91 | 100 | 90.9 | warn |
| #2 | Claude Opus 4.7 | 92.85 | 100 | 84.1 | pass |
| #3 | GPT-5.5 | 88.75 | 100 | 75 | pass |
| #4 | DeepSeek V4 Pro | 86.5 | 100 | 70 | pass |
| #5 | Qwen3 Max | 85.82 | 99 | 69.7 | pass |
| #6 | Gemini 3.1 Pro | 84.66 | 100 | 65.9 | pass |
| #7 | Grok 4 | 84.27 | 99.3 | 65.9 | pass |
| #8 | Claude Sonnet 4.6 | 69.56 | 75 | 62.9 | pass |
| #9 | Gemini 2.5 Pro | 63.95 | 69.8 | 56.8 | pass |
Data Analysis
Today's YZ Index Smoke quick test shows clear divergence among top models in how code execution and material constraint are combined. GPT-o3 has a main score of 95.91 (code execution 100, material constraint 90.9); Claude Opus 4.7 has a main score of 92.85 (code execution 100, material constraint 84.1); GPT-5.5 has a main score of 88.75 (code execution 100, material constraint 75); DeepSeek V4 Pro has a main score of 86.5 (code execution 100, material constraint 70). These models all achieve 100 or near-100 in code execution, while material constraint ranges from 90.9 to 70, forming the basis for their leading main scores.
Several models saw significant declines: Gemini 2.5 Pro fell 19.4 points on the main score, 25 points on code execution, and 12.5 points on material constraint; Claude Sonnet 4.6 fell 15.7 points on the main score and 34.9 points on material constraint; Gemini 3.1 Pro fell 11.7 points on the main score and 25.9 points on material constraint; Grok 4 fell 11 points on the main score and 23.6 points on material constraint; Qwen3 Max fell 10.2 points on the main score and 21.5 points on material constraint. DeepSeek V4 Pro plunged 19.5 points on material constraint, Qwen3 Max plunged 10.2 points on the main score, Gemini 3.1 Pro plunged 11.7 points on the main score, Grok 4 plunged 11 points on the main score, Claude Sonnet 4.6 plunged 15.7 points on the main score, and Gemini 2.5 Pro plunged 19.4 points on the main score.
The above anomalies may stem from question sampling fluctuation or may reflect genuine regression, requiring confirmation in subsequent runs. GLM-4.6 and Doubao Pro were not ranked due to incomplete data. As a small-sample single-day signal, Smoke-related observations are for reference only.
Key Changes
- Gemini 2.5 Pro: main score down 19.4 points, code execution -25, material constraint -12.5
- Claude Sonnet 4.6: main score down 15.7 points, material constraint -34.9
- Gemini 3.1 Pro: main score down 11.7 points, material constraint -25.9
- Grok 4: main score down 11 points, material constraint -23.6
- Qwen3 Max: main score down 10.2 points, material constraint -21.5
Signals to Watch
- DeepSeek V4 Pro: material constraint plunged -19.5 points
- Qwen3 Max: main score plunged -10.2 points
- Gemini 3.1 Pro: main score plunged -11.7 points
- Grok 4: main score plunged -11 points
- Claude Sonnet 4.6: main score plunged -15.7 points
- Gemini 2.5 Pro: main score plunged -19.4 points
- GLM-4.6: data incomplete (missing judgment and integrity dimensions, API failure/timeout), auto re-run initiated, not ranked this round
- Doubao Pro: data incomplete (multiple evaluation dimensions missing, API failure/timeout), auto re-run initiated, not ranked this round
When reading Smoke briefings like this, focus on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether the integrity rating has shifted from pass to warn or fail. Large single-day swings in execution or constraint scores may stem from question sampling or may be early signals of genuine regression, requiring verification in subsequent runs.
Data source: YZ Index | Run #270 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接