2026-09-28 YZ Index Smoke quick test covered 11 models, with GPT-5.5 ranking first for the day at 93.16 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals; it is not equivalent to the conclusions of the weekly Full leaderboard.
This Smoke evaluation covers only two Main Leaderboard dimensions: Code Execution and Material Constraints. The Main Leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraints. Because the daily sample size is small, single-day scores are better treated as monitoring signals rather than long-term conclusions about model capability.
Daily Ranking
| Ranking | Model | Main Leaderboard | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | GPT-5.5 | 93.16 | 95 | 90.9 | pass |
| #2 | Claude Sonnet 4.6 | 91.81 | 100 | 81.8 | pass |
| #3 | Grok 4 | 91.81 | 100 | 81.8 | pass |
| #4 | Qwen3 Max | 91.78 | 92.5 | 90.9 | warn |
| #5 | Claude Opus 4.7 | 82.16 | 75 | 90.9 | pass |
| #6 | GPT-o3 | 80.78 | 72.5 | 90.9 | pass |
| #7 | DeepSeek V4 Pro | 70.91 | 75 | 65.9 | warn |
| #8 | Gemini 3.1 Pro | 68.63 | 78.3 | 56.8 | pass |
| #9 | Doubao Pro | 68.16 | 70 | 65.9 | warn |
| #10 | Gemini 2.5 Pro | 64.06 | 70 | 56.8 | pass |
| #11 | GLM-4.6 | 51.91 | 20 | 90.9 | pass |
Data Interpretation
Among today's top three on the Main Leaderboard, GPT-5.5 scored 93.16 with a combination of Code Execution 95 and Material Constraints 90.9, while Claude Sonnet 4.6 and Grok 4 both tied at 91.81 with Code Execution 100 and Material Constraints 81.8, showing that leading models differ in their relative strengths across the two capabilities. Qwen3 Max followed closely with Code Execution 92.5, Material Constraints 90.9, and a Main Leaderboard score of 91.78, while Claude Opus 4.7 and GPT-o3 both showed a pattern of lower Code Execution with Material Constraints 90.9, with Main Leaderboard scores of 82.16 and 80.78, respectively.
On anomalous signals, Claude Opus 4.7's Main Leaderboard score plunged by 13.9 points, GPT-o3's Main Leaderboard score plunged by 15.2 points, Doubao Pro's Main Leaderboard score plunged by 14.8 points and its Integrity rating shifted from pass to warn, and GLM-4.6's Code Execution plunged by 47 points. These single-day changes may stem from question sampling fluctuations or may reflect genuine degradation, and need to be confirmed by follow-up runs using the same methodology.
The Smoke quick test has a limited sample size; the above figures are only observations for the day and should not be extrapolated to a model's overall performance.
Key Changes
- GPT-o3: Main Leaderboard down 15.2 points, Code Execution -24.5 points
- Doubao Pro: Main Leaderboard down 14.8 points, Code Execution -22 points, Material Constraints -6.1 points, Integrity pass→warn
- GPT-5.5: Main Leaderboard up 14.4 points, Code Execution +23 points
- Claude Opus 4.7: Main Leaderboard down 13.9 points, Code Execution -22 points
- Grok 4: Main Leaderboard down 5.8 points, Material Constraints -13 points
Signals to Watch
- Claude Opus 4.7: Main Leaderboard plunges by 13.9 points
- GPT-o3: Main Leaderboard plunges by 15.2 points
- Doubao Pro: Main Leaderboard plunges by 14.8 points
- GLM-4.6: Code Execution plunges by 47 points
When reading this type of Smoke brief, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether its Integrity rating has moved from pass into warn or fail. Large single-day changes in execution or constraint scores may come from question sampling or may be early signals of genuine degradation, and need to be verified by follow-up runs.
Data source: YZ Index | Run #341 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接