On 2026-09-27, the YZ Index Smoke quick test covered 10 models, with Grok 4 ranking first for the day at 97.66 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals; it is not equivalent to Full weekly ranking conclusions.
This Smoke evaluation covered only two main leaderboard dimensions: Code Execution and Material Constraints. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraints. Because the daily sample size is small, single-day scores are better used as monitoring signals rather than long-term conclusions about model capabilities.
Daily Ranking
| Rank | Model | Main Leaderboard | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | Grok 4 | 97.66 | 100 | 94.8 | pass |
| #2 | Claude Opus 4.7 | 96.01 | 97 | 94.8 | pass |
| #3 | GPT-o3 | 96.01 | 97 | 94.8 | pass |
| #4 | Claude Sonnet 4.6 | 92.19 | 93.9 | 90.1 | pass |
| #5 | Qwen3 Max | 86.44 | 79.6 | 94.8 | pass |
| #6 | Doubao Pro | 83 | 92 | 72 | pass |
| #7 | GPT-5.5 | 78.8 | 72 | 87.1 | pass |
| #8 | Gemini 2.5 Pro | 69.91 | 70 | 69.8 | pass |
| #9 | DeepSeek V4 Pro | 69.31 | 68.9 | 69.8 | pass |
| #10 | Gemini 3.1 Pro | 68.26 | 67 | 69.8 | pass |
Data Interpretation
In today's YZ Index Smoke quick test, the top models showed different patterns in the combination of Code Execution and Material Constraints. Grok 4 achieved a main leaderboard score of 97.66 with Code Execution 100 and Material Constraints 94.8, while Claude Opus 4.7 and GPT-o3 both tied at a main leaderboard score of 96.01 with Code Execution 97 and Material Constraints 94.8. Claude Sonnet 4.6 had Code Execution 93.9 and Material Constraints 90.1. Qwen3 Max had Code Execution 79.6 but Material Constraints 94.8, for a main leaderboard score of 86.44; Doubao Pro had Code Execution 92 and Material Constraints 72, for a main leaderboard score of 83, showing a combination of high Code Execution and low Material Constraints.
Compared with the previous run using the same methodology, Grok 4 rose 27.3 points on the main leaderboard, +25 points in Code Execution, and +30 points in Material Constraints; Qwen3 Max rose 11 points on the main leaderboard and +21 points in Material Constraints; GPT-5.5 fell 12.5 points on the main leaderboard and -28 points in Code Execution; Gemini 2.5 Pro fell 10.4 points on the main leaderboard and -17.5 points in Code Execution; Doubao Pro fell 9.8 points on the main leaderboard and -13.7 points in Material Constraints; and DeepSeek V4 Pro fell 8.3 points on the main leaderboard. These movements may stem from question sampling fluctuations or may reflect real changes in single-day performance, and require follow-up runs to verify.
The Smoke quick test is a small-sample single-day signal; GLM-4.6 was not ranked due to incomplete data. The score structure and changes above are drawn directly from that day's data, and no long-term inference is made for now.
Key Changes
- Grok 4: Main leaderboard up 27.3 points, Code Execution +25 points, Material Constraints +30 points
- GPT-5.5: Main leaderboard down 12.5 points, Code Execution -28 points, Material Constraints +6.4 points
- Qwen3 Max: Main leaderboard up 11 points, Material Constraints +21 points
- Gemini 2.5 Pro: Main leaderboard down 10.4 points, Code Execution -17.5 points
- Doubao Pro: Main leaderboard down 9.8 points, Code Execution -6.7 points, Material Constraints -13.7 points
Signals to Watch
- Doubao Pro: Main leaderboard plunges -9.8 points
- GPT-5.5: Main leaderboard plunges -12.5 points
- Gemini 2.5 Pro: Main leaderboard plunges -10.4 points
- DeepSeek V4 Pro: Main leaderboard plunges -8.3 points
- GLM-4.6: Incomplete data (missing communication dimension, API failure/timeout), has entered automatic rerun, not ranked in this period
When reading this type of Smoke brief, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether its integrity rating has moved from pass into warn or fail. Large single-day changes in execution or constraint scores may come from question sampling, or they may be early signals of real degradation, and require follow-up runs to verify.
Data source: YZ Index | Run #340 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接