On 2026-09-18, the YZ Index Smoke quick test covered 10 models, with Claude Opus 4.7 ranking first for the day with 100 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals and not equivalent to Full weekly ranking conclusions.
This Smoke evaluation covered only two main leaderboard dimensions: Code Execution and Material Constraints. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraints. Because the daily sample size is small, single-day scores are more suitable as monitoring signals than as long-term conclusions about model capabilities.
Today's Ranking
| Rank | Model | Main | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | Claude Opus 4.7 | 100 | 100 | 100 | pass |
| #2 | GPT-o3 | 99.01 | 100 | 97.8 | pass |
| #3 | Claude Sonnet 4.6 | 94.34 | 99.6 | 87.9 | pass |
| #4 | GPT-5.5 | 92.26 | 95 | 88.9 | pass |
| #5 | Gemini 2.5 Pro | 90.91 | 100 | 79.8 | pass |
| #6 | Doubao Pro | 83.5 | 70 | 100 | warn |
| #7 | Qwen3 Max | 82.42 | 89.3 | 74 | pass |
| #8 | Grok 4 | 81.17 | 75 | 88.7 | pass |
| #9 | Gemini 3.1 Pro | 69.75 | 45 | 100 | pass |
| #10 | DeepSeek V4 Pro | 67.51 | 50 | 88.9 | pass |
Data Interpretation
In today's YZ Index Smoke quick test, Claude Opus 4.7 ranked first with a structure of Main 100, Code Execution 100, and Material Constraints 100, forming a completely balanced combination of Code Execution and Material Constraints. GPT-o3 had Main 99.01, Code Execution 100, and Material Constraints 97.8, showing that Code Execution remained at a full score while Material Constraints were slightly lower. Claude Sonnet 4.6 had Main 94.34, Code Execution 99.6, and Material Constraints 87.9, likewise showing a combination in which Code Execution is stronger than Material Constraints. Gemini 2.5 Pro had Main 90.91, Code Execution 100, and Material Constraints 79.8, further highlighting the characteristic of high Code Execution and low Material Constraints. By contrast, Doubao Pro had Main 83.5, Code Execution 70, and Material Constraints 100, reflecting a complementary structure with a full score in Material Constraints but relatively weak Code Execution.
In terms of significant changes, Claude Opus 4.7's Main score rose 33.2 points compared with the previous same-basis run, with Code Execution up 53 points and Material Constraints up 9.1 points; Qwen3 Max's Main score rose 27.8 points, with Code Execution up 64.3 points but Material Constraints down 16.9 points; GPT-o3's Main score rose 21.6 points, with Code Execution up 28 points and Material Constraints up 13.7 points. These increases were mainly driven by the Code Execution dimension, while Material Constraints diverged. Qwen3 Max's 16.9-point plunge in Material Constraints may stem from single-day question sampling fluctuation, and may also represent real degradation; subsequent runs are needed to verify signal stability. GLM-4.6 did not participate in this period's ranking because an API failure led to incomplete data.
Overall, leading models have different emphases in the strength mix between Code Execution and Material Constraints; score changes among the models with notable moves are concentrated in the Code Execution dimension, while fluctuations on the Material Constraints side are relatively limited. As a small-sample single-day signal, Smoke's observations above only reflect that day's data characteristics and do not constitute a basis for long-term judgment.
Main Changes
- Claude Opus 4.7: Main up 33.2 points, Code Execution +53 points, Material Constraints +9.1 points
- Qwen3 Max: Main up 27.8 points, Code Execution +64.3 points, Material Constraints -16.9 points
- GPT-o3: Main up 21.6 points, Code Execution +28 points, Material Constraints +13.7 points
- DeepSeek V4 Pro: Main up 19.7 points, Material Constraints +43.7 points
- Claude Sonnet 4.6: Main up 17.9 points, Code Execution +27.6 points, Material Constraints +6.1 points
Signals to Watch
- Qwen3 Max: Material Constraints plunge of -16.9 points
- GLM-4.6: Incomplete data (missing execution, judgment, integrity, and communication dimensions; API failure/timeout); an automatic rerun has been initiated, and it does not participate in this period's ranking
When reading this kind of Smoke brief, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether its Integrity rating has moved from pass to warn or fail. Large single-day changes in execution or constraint scores may come from question sampling, or may be an early signal of real degradation, requiring subsequent runs for review.
Data source: YZ Index | Run #328 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接