The 2026-09-21 YZ Index Smoke quick test covers 10 models, with Grok 4 ranking first for the day at 95.44 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals and not equivalent to Full weekly leaderboard conclusions.
This Smoke evaluation covers only two main leaderboard dimensions: Code Execution and Material Constraints. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraints. Because the daily sample size is small, a single-day score is better used as a monitoring signal rather than a long-term verdict on model capability.
Daily Ranking
| Rank | Model | Main Leaderboard | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | Grok 4 | 95.44 | 91.7 | 100 | warn |
| #2 | Gemini 3.1 Pro | 95.19 | 100 | 89.3 | pass |
| #3 | DeepSeek V4 Pro | 95.01 | 100 | 88.9 | pass |
| #4 | Doubao Pro | 90.62 | 91.7 | 89.3 | pass |
| #5 | Gemini 2.5 Pro | 85.63 | 91.7 | 78.2 | pass |
| #6 | GPT-o3 | 81.44 | 75 | 89.3 | pass |
| #7 | GPT-5.5 | 76.44 | 75 | 78.2 | pass |
| #8 | Claude Sonnet 4.6 | 72.5 | 50 | 100 | pass |
| #9 | Qwen3 Max | 72.25 | 58.3 | 89.3 | pass |
| #10 | Claude Opus 4.7 | 70.77 | 55.6 | 89.3 | pass |
Data Interpretation
In today's YZ Index Smoke quick test, the leading models show different characteristics in the combination of Code Execution and Material Constraints. Grok 4 has a main leaderboard score of 95.44, Code Execution of 91.7, and Material Constraints of 100, with its Material Constraints score standing out. Gemini 3.1 Pro has a main leaderboard score of 95.19, Code Execution of 100, and Material Constraints of 89.3, reaching a perfect score in the Code Execution dimension. DeepSeek V4 Pro has a main leaderboard score of 95.01, Code Execution of 100, and Material Constraints of 88.9, also supported by a high Code Execution score. Doubao Pro has a main leaderboard score of 90.62, Code Execution of 91.7, and Material Constraints of 89.3, with the two dimensions relatively balanced.
Compared with the previous run using the same methodology, some models showed significant fluctuations. Claude Opus 4.7 fell 26.2 points on the main leaderboard and 44.4 points in Code Execution. Grok 4 rose 20.1 points on the main leaderboard, 17.4 points in Code Execution, and 23.3 points in Material Constraints. Gemini 3.1 Pro rose 19.8 points on the main leaderboard, 25.7 points in Code Execution, and 12.6 points in Material Constraints. DeepSeek V4 Pro rose 19.2 points on the main leaderboard, 25 points in Code Execution, and 12.2 points in Material Constraints. Gemini 2.5 Pro rose 18.4 points on the main leaderboard, 41.7 points in Code Execution, and fell 10.1 points in Material Constraints.
In terms of anomalous signals, GPT-5.5 plunged 10.1 points on the main leaderboard, Claude Sonnet 4.6 plunged 17 points, Qwen3 Max plunged 17.3 points, and Claude Opus 4.7 plunged 26.2 points. These changes may stem from question sampling variance or genuine degradation and need to be confirmed by subsequent runs. GLM-4.6 has incomplete data, has entered automatic rerun, and is not included in this period's ranking. As Smoke is a small-sample single-day signal, the above observations are for reference only.
Major Changes
- Claude Opus 4.7: Fell 26.2 points on the main leaderboard and 44.4 points in Code Execution
- Grok 4: Rose 20.1 points on the main leaderboard, 17.4 points in Code Execution, and 23.3 points in Material Constraints; Integrity pass→warn
- Gemini 3.1 Pro: Rose 19.8 points on the main leaderboard, 25.7 points in Code Execution, and 12.6 points in Material Constraints
- DeepSeek V4 Pro: Rose 19.2 points on the main leaderboard, 25 points in Code Execution, and 12.2 points in Material Constraints
- Gemini 2.5 Pro: Rose 18.4 points on the main leaderboard, 41.7 points in Code Execution, and fell 10.1 points in Material Constraints
Signals to Watch
- GPT-5.5: Main leaderboard plunged 10.1 points
- Claude Sonnet 4.6: Main leaderboard plunged 17 points
- Qwen3 Max: Main leaderboard plunged 17.3 points
- Claude Opus 4.7: Main leaderboard plunged 26.2 points
- GLM-4.6: Incomplete data (missing execution, judgment, integrity, and communication dimensions; API failure/timeout), has entered automatic rerun, and is not included in this period's ranking
When reading this kind of Smoke brief, the focus should be on two questions: first, whether a model has repeatedly exposed the same type of weakness over multiple days; second, whether its Integrity rating has moved from pass to warn or fail. Large single-day changes in Code Execution or Material Constraints scores may come from question sampling or may be early signals of genuine degradation, requiring confirmation in subsequent runs.
Data source: YZ Index | Run #332 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接