On 2026-09-11, the YZ Index Smoke quick test covered 10 models, with Doubao Pro ranking first that day at 95.28. Smoke is a daily 10-question quick test, suitable for observing short-term signals, and is not equivalent to Full weekly leaderboard conclusions.
This Smoke evaluation covers only two main leaderboard dimensions: code execution and material constraints. The main leaderboard formula is 0.55 × code execution + 0.45 × material constraints. Because the daily sample size is small, single-day scores are better used as monitoring signals rather than as long-term conclusions about model capability.
Daily Ranking
| Ranking | Model | Main leaderboard | Code execution | Material constraints | Integrity |
|---|---|---|---|---|---|
| #1 | Doubao Pro | 95.28 | 100 | 89.5 | pass |
| #2 | Gemini 3.1 Pro | 93.48 | 100 | 85.5 | pass |
| #3 | GPT-o3 | 93.48 | 100 | 85.5 | pass |
| #4 | GPT-5.5 | 90.82 | 83.3 | 100 | pass |
| #5 | Grok 4 | 88.64 | 97.5 | 77.8 | pass |
| #6 | DeepSeek V4 Pro | 86.09 | 83.3 | 89.5 | pass |
| #7 | Gemini 2.5 Pro | 81.53 | 75 | 89.5 | pass |
| #8 | Claude Opus 4.7 | 79.1 | 75 | 84.1 | pass |
| #9 | Claude Sonnet 4.6 | 74.89 | 72.5 | 77.8 | pass |
| #10 | Qwen3 Max | 70.91 | 75 | 65.9 | pass |
Data Interpretation
In today's YZ Index Smoke quick test, Doubao Pro ranked first on the main leaderboard with 95.28; its combination of 100 in code execution and 89.5 in material constraints shows a balanced advantage. Gemini 3.1 Pro and GPT-o3 tied on the main leaderboard at 93.48, with both scoring 100 in code execution and 85.5 in material constraints, reflecting a structural emphasis on code execution. GPT-5.5's main leaderboard score of 90.82 relied on 100 in material constraints and 83.3 in code execution, forming a complementary profile. Leading models show clear divergence in how they combine strengths and weaknesses across the two dimensions.
Compared with the previous comparable run, GPT-o3 rose 15.4 points on the main leaderboard and 30.2 points in code execution; Gemini 2.5 Pro fell 7.2 points on the main leaderboard, fell 25 points in code execution, and rose 14.5 points in material constraints; GPT-5.5 rose 6.2 points on the main leaderboard and 11.3 points in code execution; Grok 4 fell 5.2 points on the main leaderboard, rose 8.8 points in code execution, and fell 22.2 points in material constraints; Claude Opus 4.7 rose 5.2 points on the main leaderboard, rose 22.4 points in code execution, and fell 15.9 points in material constraints. Among these shifts, the sharp drop of -22.2 points in Grok 4's material constraints, the sharp drop of -25 points in Gemini 2.5 Pro's code execution, and the sharp drop of -15.9 points in Claude Opus 4.7's material constraints constitute anomalous signals. They may stem from question-sampling volatility or may indicate genuine degradation, and need follow-up runs to verify.
GLM-4.6 had incomplete data due to API failure/timeout and did not participate in this period's ranking. As Smoke is a small-sample, single-day signal, the above observations reflect only that day's score structure and should not be used for long-term inferences for now.
Main Changes
- GPT-o3: Main leaderboard up 15.4 points, code execution +30.2 points
- Gemini 2.5 Pro: Main leaderboard down 7.2 points, code execution -25 points, material constraints +14.5 points
- GPT-5.5: Main leaderboard up 6.2 points, code execution +11.3 points
- Grok 4: Main leaderboard down 5.2 points, code execution +8.8 points, material constraints -22.2 points
- Claude Opus 4.7: Main leaderboard up 5.2 points, code execution +22.4 points, material constraints -15.9 points
Signals to Watch
- Grok 4: Material constraints plunge -22.2 points
- Gemini 2.5 Pro: Code execution plunge -25 points
- Claude Opus 4.7: Material constraints plunge -15.9 points
- GLM-4.6: Incomplete data (missing several evaluation dimensions, including execution, judgment, integrity, and communication; API failure/timeout), has entered automatic rerun, and is not included in this period's ranking
When reading this kind of Smoke brief, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether its integrity rating moves from pass to warn or fail. Large single-day changes in execution or constraint scores may come from question sampling or may be early signals of genuine degradation, requiring follow-up runs to verify.
Data source: YZ Index | Run #318 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接