The 2026-10-04 YZ Index Smoke quick test covered 15 models, with Claude Opus 4.7 leading the day at 81.57 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals, and is not equivalent to Full weekly leaderboard conclusions.
This Smoke evaluation covers only two main leaderboard dimensions: Code Execution and Material Constraints. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraints. Because the daily sample size is small, a single-day score is better suited as a monitoring signal rather than a long-term conclusion about model capability.
Daily Ranking
| Rank | Model | Main Leaderboard | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | Claude Opus 4.7 | 81.57 | 97 | 62.7 | pass |
| #2 | GPT-o3 | 79.44 | 97.8 | 57 | pass |
| #3 | Doubao Pro | 75.57 | 96 | 50.6 | pass |
| #4 | Grok 4 | 74.2 | 95.3 | 48.4 | pass |
| #5 | DeepSeek V4 Pro | 73.52 | 75 | 71.7 | pass |
| #6 | Gemini 3.1 Pro | 72.53 | 75 | 69.5 | pass |
| #7 | Qwen3 Max | 66.69 | 61.6 | 72.9 | pass |
| #8 | GPT-6 Luna | 64.27 | 69.8 | 57.5 | pass |
| #9 | Gemini 2.5 Pro | 63.68 | 69.8 | 56.2 | pass |
| #10 | GPT-6.1 Sol | 62.51 | 69.8 | 53.6 | pass |
| #11 | Claude Sonnet 4.6 | 62.32 | 71.9 | 50.6 | pass |
| #12 | GPT-6 Sol | 61.16 | 69.8 | 50.6 | pass |
| #13 | GPT-5.5 | 60.11 | 75 | 41.9 | pass |
| #14 | GPT-6 Astra | 52.86 | 44.8 | 62.7 | pass |
| #15 | GLM-4.6 | 33.94 | 50 | 14.3 | warn |
Data Interpretation
Among today's top three on the main leaderboard, Claude Opus 4.7 scored 81.57 with a combination of Code Execution 97 and Material Constraints 62.7; GPT-o3 reached 79.44 by relying on Code Execution 97.8 and Material Constraints 57; Doubao Pro ranked third with Code Execution 96 and Material Constraints 50.6. These models remain high on the Code Execution dimension while scoring relatively low on Material Constraints, creating a clearly imbalanced profile. DeepSeek V4 Pro and Gemini 3.1 Pro show more balanced combinations: the former has Code Execution 75 and Material Constraints 71.7, while the latter has Code Execution 75 and Material Constraints 69.5, with main leaderboard scores of 73.52 and 72.53, respectively.
Several models showed notable declines versus the previous same-basis run: GPT-6 Sol fell 24.1 points on the main leaderboard and 47.2 points on Material Constraints; Gemini 2.5 Pro fell 23.7 points on the main leaderboard and 30.2 points on Code Execution; GPT-6 Astra fell 21.9 points on the main leaderboard and 32.1 points on Material Constraints; Material Constraints for Claude Sonnet 4.6 and Claude Opus 4.7 fell 36.2 and 34.3 points, respectively. Smoke tests are small-sample, single-day signals; such changes may stem from question sampling fluctuations or may reflect genuine capability degradation, and require follow-up runs to confirm.
Overall, leading models mostly rely on Code Execution advantages to lift their main leaderboard scores, while combinations with weaker Material Constraints are relatively concentrated in today's data. Qwen3 Max's inverse structure—Code Execution 61.6 and Material Constraints 72.9—gives it a main leaderboard score of 66.69, showing the direct impact of different weighted combinations on ranking. All interpretations are based on exact values for the day and no extrapolation has been made.
Key Changes
- GPT-6 Sol: Main leaderboard down 24.1 points, Code Execution -5.2 points, Material Constraints -47.2 points
- Gemini 2.5 Pro: Main leaderboard down 23.7 points, Code Execution -30.2 points, Material Constraints -15.8 points
- GPT-6 Astra: Main leaderboard down 21.9 points, Code Execution -13.5 points, Material Constraints -32.1 points
- Claude Sonnet 4.6: Main leaderboard down 17.3 points, Material Constraints -36.2 points
- Claude Opus 4.7: Main leaderboard down 17.1 points, Material Constraints -34.3 points
Signals to Watch
- No publishable anomaly signals were retained for this run.
When reading this type of Smoke brief, focus on two questions: first, whether a model exposes the same type of weakness on multiple consecutive days; second, whether its Integrity rating moves from pass to warn or fail. Large single-day changes in execution or constraint scores may come from question sampling or may be early signals of genuine degradation, requiring follow-up runs to verify.
Data source: YZ Index | Run #359 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接