The 2026-09-13 YZ Index Smoke quick test covered 10 models, with Grok 4 ranking first for the day at 87 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals; it is not equivalent to the conclusions of the weekly Full leaderboard.
This Smoke evaluation covered only two main leaderboard dimensions: Code Execution and Material Constraints. The Main Leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraints. Because the daily sample size is small, single-day scores are better used as monitoring signals rather than long-term conclusions about model capabilities.
Daily Ranking
| Rank | Model | Main Leaderboard | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | Grok 4 | 87 | 100 | 71.1 | pass |
| #2 | Doubao Pro | 85.15 | 100 | 67 | pass |
| #3 | GPT-o3 | 85.15 | 100 | 67 | pass |
| #4 | DeepSeek V4 Pro | 77.34 | 75 | 80.2 | pass |
| #5 | Gemini 3.1 Pro | 73.25 | 75 | 71.1 | pass |
| #6 | Claude Opus 4.7 | 63.59 | 50 | 80.2 | pass |
| #7 | GPT-5.5 | 60.87 | 75 | 43.6 | pass |
| #8 | Gemini 2.5 Pro | 59.5 | 50 | 71.1 | pass |
| #9 | Qwen3 Max | 57.65 | 50 | 67 | pass |
| #10 | Claude Sonnet 4.6 | 53.56 | 50 | 57.9 | pass |
Data Interpretation
Today's YZ Index Smoke quick test shows clear divergence among leading models in how Code Execution and Material Constraints pair together. Grok 4 has a Main Leaderboard score of 87, with Code Execution 100 and Material Constraints 71.1; Doubao Pro and GPT-o3 both have Main Leaderboard scores of 85.15, both with Code Execution 100 and Material Constraints 67. All three occupy top positions on the strength of high Code Execution scores. DeepSeek V4 Pro has a Main Leaderboard score of 77.34, with Code Execution 75 and Material Constraints 80.2; Gemini 3.1 Pro has a Main Leaderboard score of 73.25, with Code Execution 75 and Material Constraints 71.1. The two have relatively stronger Material Constraints, forming a different balance structure from the top three.
Several models showed significant changes: Qwen3 Max fell 25.9 points on the Main Leaderboard and 49.2 points in Code Execution; Claude Sonnet 4.6 fell 22.6 points on the Main Leaderboard and 45.5 points in Code Execution, while Material Constraints rose 5.3 points; GPT-5.5 fell 21.4 points on the Main Leaderboard, 22 points in Code Execution, and 20.7 points in Material Constraints; Claude Opus 4.7 fell 18.7 points on the Main Leaderboard and 47 points in Code Execution, while Material Constraints rose 15.9 points; Gemini 2.5 Pro fell 16.6 points on the Main Leaderboard, 23.5 points in Code Execution, and 8.2 points in Material Constraints. These anomalous signals may stem from question sampling fluctuations, or they may reflect genuine single-day performance regression, and require follow-up run review for confirmation.
GLM-4.6's data was incomplete due to API failures/timeouts, with several missing evaluation dimensions, so it did not participate in this edition's ranking. As a small-sample, single-day signal, the observations above only reflect the day's score structure, and the wording remains measured.
Key Changes
- Qwen3 Max: Main Leaderboard down 25.9 points, Code Execution -49.2 points
- Claude Sonnet 4.6: Main Leaderboard down 22.6 points, Code Execution -45.5 points, Material Constraints +5.3 points
- GPT-5.5: Main Leaderboard down 21.4 points, Code Execution -22 points, Material Constraints -20.7 points
- Claude Opus 4.7: Main Leaderboard down 18.7 points, Code Execution -47 points, Material Constraints +15.9 points
- Gemini 2.5 Pro: Main Leaderboard down 16.6 points, Code Execution -23.5 points, Material Constraints -8.2 points
Signals to Watch
- Claude Opus 4.7: Main Leaderboard plunged -18.7 points
- GPT-5.5: Main Leaderboard plunged -21.4 points
- Gemini 2.5 Pro: Main Leaderboard plunged -16.6 points
- Qwen3 Max: Main Leaderboard plunged -25.9 points
- Claude Sonnet 4.6: Main Leaderboard plunged -22.6 points
- GLM-4.6: Data incomplete (missing multiple evaluation dimensions due to API failure/timeout); has been queued for automatic rerun and is not included in this edition's ranking
When reading this type of Smoke brief, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether its Integrity rating has moved from pass to warn or fail. Large single-day changes in Code Execution or Material Constraints scores may come from question sampling, or they may be early signals of genuine regression, requiring follow-up run review.
Data: YZ Index | Run #320 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接