The 2026-09-14 YZ Index Smoke quick test covered 10 models, with Claude Opus 4.7, DeepSeek V4 Pro, Doubao Pro, and GPT-5.5 tying for first place that day at 91.09. Smoke is a daily 10-question quick test, suited to observing short-term signals; it is not equivalent to Full weekly leaderboard conclusions.
This Smoke evaluation covers only two main leaderboard dimensions: code execution and material constraints. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraints. Because the daily sample size is small, single-day scores are better used as monitoring signals rather than as long-term conclusions about model capability.
Daily Ranking
| Rank | Model | Main Leaderboard | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | Claude Opus 4.7 | 91.09 | 100 | 80.2 | pass |
| #2 | DeepSeek V4 Pro | 91.09 | 100 | 80.2 | pass |
| #3 | Doubao Pro | 91.09 | 100 | 80.2 | pass |
| #4 | GPT-5.5 | 91.09 | 100 | 80.2 | pass |
| #5 | Gemini 3.1 Pro | 87 | 100 | 71.1 | pass |
| #6 | Grok 4 | 86.61 | 99.3 | 71.1 | pass |
| #7 | Claude Sonnet 4.6 | 82.5 | 100 | 61.1 | pass |
| #8 | Qwen3 Max | 70.47 | 62.5 | 80.2 | warn |
| #9 | Gemini 2.5 Pro | 70.19 | 75 | 64.3 | pass |
| #10 | GPT-o3 | 70.19 | 75 | 64.3 | pass |
Data Interpretation
The top four models on today's main leaderboard—Claude Opus 4.7, DeepSeek V4 Pro, Doubao Pro, and GPT-5.5—all show a structure of 100 in code execution and 80.2 in material constraints, with the same main leaderboard score of 91.09, indicating a balanced mix of strengths and weaknesses with consistent values. Gemini 3.1 Pro also scores 100 in code execution but drops to 71.1 in material constraints, with its main leaderboard score falling to 87; Grok 4 scores 99.3 in code execution and 71.1 in material constraints, for a main leaderboard score of 86.61; Claude Sonnet 4.6 scores 100 in code execution and 61.1 in material constraints, for a main leaderboard score of 82.5, reflecting the impact of the material constraints gap on overall placement.
GPT-5.5's main leaderboard score rose 30.2 points from the previous same-scope run, with code execution up 25 points and material constraints up 36.6 points; Claude Sonnet 4.6's main leaderboard score rose 28.9 points, with code execution up 50 points; Claude Opus 4.7's main leaderboard score rose 27.5 points, with code execution up 50 points; DeepSeek V4 Pro's main leaderboard score rose 13.8 points, with code execution up 25 points; GPT-o3's main leaderboard score fell 15 points, with code execution down 25 points. These changes all come from single-day Smoke small-sample data and require follow-up runs to distinguish question-sampling fluctuations from genuine degradation.
GPT-o3 shows an anomalous -15-point signal on the main leaderboard, which may stem from single-day sampling fluctuation or a temporary change in model performance and requires a full run for verification; GLM-4.6 did not participate in the ranking because its data was incomplete (missing several evaluation dimensions; API failure or timeout). An automatic rerun has been triggered, and this period's results are for daily reference only.
Key Changes
- GPT-5.5: main leaderboard up 30.2 points, code execution +25 points, material constraints +36.6 points
- Claude Sonnet 4.6: main leaderboard up 28.9 points, code execution +50 points
- Claude Opus 4.7: main leaderboard up 27.5 points, code execution +50 points
- GPT-o3: main leaderboard down 15 points, code execution -25 points
- DeepSeek V4 Pro: main leaderboard up 13.8 points, code execution +25 points
Signals to Watch
- GPT-o3: main leaderboard showed a sharp -15-point drop
- GLM-4.6: incomplete data (missing several evaluation dimensions; API failure/timeout), has entered automatic rerun, and does not participate in this period's ranking
When reading this kind of Smoke brief, focus on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether its integrity rating has moved from pass to warn or fail. Large single-day changes in execution or constraint scores may come from question sampling or may be early signals of genuine degradation, requiring follow-up runs for verification.
Data Source: YZ Index | Run #322 | View Raw Data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接