The 2026-10-11 YZ Index Smoke quick test covered 13 models, with Claude Opus 4.7, GPT-6 Astra, and GPT-6.1 Sol tied for first place that day at 79.79 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals; it is not equivalent to the conclusions of the weekly Full leaderboard.
This Smoke evaluation covered only two main leaderboard dimensions: code execution and material constraints. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraints. Because the daily sample size is small, single-day scores are better treated as monitoring signals rather than as long-term verdicts on model capability.
Daily Rankings
| Rank | Model | Main Leaderboard | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | Claude Opus 4.7 | 79.79 | 72 | 89.3 | pass |
| #2 | GPT-6 Astra | 79.79 | 72 | 89.3 | pass |
| #3 | GPT-6.1 Sol | 79.79 | 72 | 89.3 | pass |
| #4 | Claude Sonnet 4.6 | 74.52 | 72 | 77.6 | pass |
| #5 | Gemini 3.1 Pro | 68.54 | 72 | 64.3 | pass |
| #6 | Gemini 2.5 Pro | 66.04 | 47 | 89.3 | pass |
| #7 | GPT-6 Sol | 66.04 | 47 | 89.3 | pass |
| #8 | Grok 4 | 64.92 | 75 | 52.6 | pass |
| #9 | GPT-6 Luna | 62.1 | 72 | 50 | pass |
| #10 | Doubao Pro | 54.79 | 47 | 64.3 | pass |
| #11 | GPT-o3 | 54.79 | 47 | 64.3 | pass |
| #12 | DeepSeek V4 Pro | 49.52 | 47 | 52.6 | pass |
| #13 | GPT-5.5 | 48.35 | 47 | 50 | pass |
Data Interpretation
Today's top three on the main leaderboard — Claude Opus 4.7, GPT-6 Astra, and GPT-6.1 Sol — all scored 79.79, with code execution at 72 and material constraints at 89.3. The three are structurally identical, showing that while they maintain a high level on the material constraints dimension, their code execution scores fall in the same range. Fourth-place Claude Sonnet 4.6 followed at 74.52, with code execution also at 72 but material constraints dropping to 77.6, reflecting the impact of relatively weaker material constraints on the overall ranking. Gemini 3.1 Pro recorded 72 for code execution and 64.3 for material constraints, for a main leaderboard score of 68.54, while Grok 4 scored 64.92 with 75 for code execution and 52.6 for material constraints — illustrating the different ways strengths in code execution and material constraints combined across models in the day's score distribution.
DeepSeek V4 Pro's main leaderboard score fell by 47.4 points, with code execution down 53 points and material constraints down 40.6 points; GPT-5.5 fell 38.6 points on the main leaderboard, with code execution down 53 points and material constraints down 20.9 points; GPT-o3 fell 35 points on the main leaderboard, with code execution down 51.5 points and material constraints down 14.8 points; Doubao Pro fell 32.1 points on the main leaderboard, with code execution down 53 points and material constraints down 6.6 points; GPT-6 Luna fell 26.7 points on the main leaderboard, with code execution down 28 points and material constraints down 25 points. These significant changes may stem from sampling fluctuations in the day's questions, or they may reflect genuine performance degradation, and require follow-up runs using the same methodology to confirm. The Smoke quick test is a small-sample, single-day signal; the observations above are for reference that day only and do not constitute a basis for long-term judgment.
Key Changes
- DeepSeek V4 Pro: main leaderboard down 47.4 points, code execution −53 points, material constraints −40.6 points
- GPT-5.5: main leaderboard down 38.6 points, code execution −53 points, material constraints −20.9 points
- GPT-o3: main leaderboard down 35 points, code execution −51.5 points, material constraints −14.8 points
- Doubao Pro: main leaderboard down 32.1 points, code execution −53 points, material constraints −6.6 points
- GPT-6 Luna: main leaderboard down 26.7 points, code execution −28 points, material constraints −25 points
Signals to Watch
- No publishable anomalous signals were retained in this run.
When reading this kind of Smoke brief, the focus should be on two questions: first, whether a given model shows the same type of weakness on consecutive days; second, whether its integrity rating moves from pass into warn or fail. Large single-day changes in execution or constraint scores may come from question sampling, or they may be early signals of genuine degradation, and require follow-up runs to confirm.
Data source: YZ Index | Run #371 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接