On 2026-10-10, the YZ Index Smoke quick test covered 15 models, and DeepSeek V4 Pro ranked first for the day with 96.94 points. Smoke is a daily 10-question quick test suitable for observing short-term signals; it is not equivalent to Full weekly leaderboard conclusions.
This Smoke evaluation covered only two main-leaderboard dimensions: code execution and material constraints. The main leaderboard formula is 0.55 × code execution + 0.45 × material constraints. Because the daily sample size is small, single-day scores are better used as monitoring signals than as long-term conclusions about model capability.
Daily Rankings
| Ranking | Model | Main | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | DeepSeek V4 Pro | 96.94 | 100 | 93.2 | pass |
| #2 | GPT-6.1 Sol | 92.85 | 100 | 84.1 | pass |
| #3 | Claude Opus 4.7 | 92.02 | 98.5 | 84.1 | pass |
| #4 | Gemini 3.1 Pro | 91.81 | 100 | 81.8 | pass |
| #5 | GPT-o3 | 89.77 | 98.5 | 79.1 | pass |
| #6 | GPT-6 Astra | 88.75 | 100 | 75 | pass |
| #7 | GPT-6 Luna | 88.75 | 100 | 75 | pass |
| #8 | Doubao Pro | 86.91 | 100 | 70.9 | pass |
| #9 | GPT-5.5 | 86.91 | 100 | 70.9 | pass |
| #10 | Claude Sonnet 4.6 | 82.41 | 100 | 60.9 | pass |
| #11 | GPT-6 Sol | 77.59 | 100 | 50.2 | pass |
| #12 | Qwen3 Max | 76.75 | 87.5 | 63.6 | pass |
| #13 | Grok 4 | 75 | 75 | 75 | warn |
| #14 | Gemini 2.5 Pro | 74.18 | 73.5 | 75 | pass |
| #15 | GLM-4.6 | 45.66 | 25 | 70.9 | pass |
Data Interpretation
In today's YZ Index Smoke quick test, DeepSeek V4 Pro ranked first on the main leaderboard with 96.94, with a structure of code execution 100 and material constraints 93.2 showing a standout balance. GPT-6.1 Sol had code execution 100 and material constraints 84.1, for a main leaderboard score of 92.85; Gemini 3.1 Pro had code execution 100 and material constraints 81.8, for a main leaderboard score of 91.81; Claude Opus 4.7 had code execution 98.5 and material constraints 84.1, for a main leaderboard score of 92.02. Leading models generally maintained high code execution, while differences in material constraints directly affected main leaderboard rankings.
GLM-4.6 scored 45.66 on the main leaderboard, with code execution 25 and material constraints 70.9, down 31.4 points on the main leaderboard, down 33.3 points in code execution, and down 29.1 points in material constraints from the previous run. GPT-o3 scored 89.77 on the main leaderboard, with code execution 98.5 and material constraints 79.1, up 18 points on the main leaderboard and up 40.2 points in code execution versus the previous run. Gemini 3.1 Pro rose 16.8 points on the main leaderboard and 25 points in code execution. Claude Opus 4.7 saw material constraints plunge 15.9 points, GPT-6 Luna saw material constraints plunge 20 points, Doubao Pro saw material constraints plunge 29.1 points, Claude Sonnet 4.6 fell 12.3 points on the main leaderboard, and GPT-6 Sol saw material constraints plunge 19.6 points. These anomalous signals may come from question sampling fluctuations or genuine degradation and require follow-up runs for confirmation. The Smoke quick test is a small-sample, single-day signal, and wording should remain restrained.
Key Changes
- GLM-4.6: Main leaderboard down 31.4 points, code execution -33.3 points, material constraints -29.1 points, integrity warn→pass
- GPT-o3: Main leaderboard up 18 points, code execution +40.2 points, material constraints -9.2 points
- Gemini 3.1 Pro: Main leaderboard up 16.8 points, code execution +25 points, material constraints +6.8 points
- GPT-6 Luna: Main leaderboard up 13.9 points, code execution +41.7 points, material constraints -20 points
- Claude Sonnet 4.6: Main leaderboard down 12.3 points, material constraints -27.4 points
Signals to Watch
- Claude Opus 4.7: Material constraints plunge -15.9 points
- GPT-6 Luna: Material constraints plunge -20 points
- Doubao Pro: Material constraints plunge -29.1 points
- Claude Sonnet 4.6: Main leaderboard plunge -12.3 points
- GPT-6 Sol: Material constraints plunge -19.6 points
- GLM-4.6: Main leaderboard plunge -31.4 points
When reading this type of Smoke brief, the focus should be on two questions: first, whether a model has repeatedly shown the same type of weakness over multiple consecutive days; second, whether its integrity rating has moved from pass into warn or fail. Large single-day changes in execution or constraint scores may come from question sampling, or they may be early signals of genuine degradation, requiring follow-up runs.
Data source: YZ Index | Run #370 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接