On 2026-09-24, the YZ Index Smoke quick test covered 10 models, with Claude Opus 4.7 and Gemini 3.1 Pro tied for first place for the day at 84.61. Smoke is a daily 10-question quick test, suitable for observing short-term signals, and is not equivalent to Full weekly ranking conclusions.
This Smoke evaluation covered only the two main leaderboard dimensions of Code Execution and Material Constraints. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraints. Because the daily sample size is small, single-day scores are better used as monitoring signals rather than long-term conclusions about model capabilities.
Daily Rankings
| Rank | Model | Main Leaderboard | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | Claude Opus 4.7 | 84.61 | 100 | 65.8 | pass |
| #2 | Gemini 3.1 Pro | 84.61 | 100 | 65.8 | pass |
| #3 | Grok 4 | 82.35 | 95.8 | 65.9 | pass |
| #4 | Doubao Pro | 82.3 | 95.8 | 65.8 | pass |
| #5 | DeepSeek V4 Pro | 79.35 | 100 | 54.1 | pass |
| #6 | Gemini 2.5 Pro | 78.21 | 95.8 | 56.7 | pass |
| #7 | GPT-o3 | 76.33 | 100 | 47.4 | pass |
| #8 | GPT-5.5 | 70.86 | 75 | 65.8 | pass |
| #9 | Qwen3 Max | 69.83 | 75 | 63.5 | pass |
| #10 | Claude Sonnet 4.6 | 51.85 | 50 | 54.1 | pass |
Data Interpretation
In today's YZ Index Smoke quick test, Claude Opus 4.7 and Gemini 3.1 Pro tied at 84.61 on the main leaderboard, both showing a combination of 100 in Code Execution and 65.8 in Material Constraints, indicating that both reached a perfect score in the Code Execution dimension while maintaining a mid-to-high level in Material Constraints. Grok 4 and Doubao Pro followed closely, with main leaderboard scores of 82.35 and 82.3 respectively; both scored 95.8 in Code Execution, and their Material Constraints scores were 65.9 and 65.8 respectively, reflecting a relatively balanced structure between Code Execution and Material Constraints. DeepSeek V4 Pro also scored 100 in Code Execution, but its Material Constraints score was only 54.1, with a main leaderboard score of 79.35, reflecting its relatively weaker characteristic in the Material Constraints dimension.
Compared with the previous run under the same methodology, Claude Sonnet 4.6's main leaderboard score dropped 45.1 points, Code Execution dropped 44.5 points, and Material Constraints dropped 45.9 points; Qwen3 Max's main leaderboard score dropped 21.4 points, Code Execution dropped 10.9 points, and Material Constraints dropped 34.3 points; GPT-5.5's main leaderboard score dropped 21 points, Code Execution dropped 14.5 points, and Material Constraints dropped 29 points. These changes may stem from question sampling fluctuations or may reflect real single-day performance differences, and follow-up runs are needed to confirm stability. Gemini 2.5 Pro's Material Constraints dropped 40.3 points, and Claude Opus 4.7's Material Constraints dropped 34.2 points, also suggesting the need for further observation.
The Smoke quick test is a small-sample, single-day signal. The current data only shows immediate differences among leading models in the pairing of Code Execution and Material Constraints. The specific reasons for the anomalous models still await multiple rounds of confirmation.
Main Changes
- Claude Sonnet 4.6: Main leaderboard down 45.1 points, Code Execution -44.5, Material Constraints -45.9
- Qwen3 Max: Main leaderboard down 21.4 points, Code Execution -10.9, Material Constraints -34.3
- GPT-5.5: Main leaderboard down 21 points, Code Execution -14.5, Material Constraints -29
- Gemini 2.5 Pro: Main leaderboard down 17 points, Material Constraints -40.3
- Claude Opus 4.7: Main leaderboard down 12.2 points, Code Execution +5.9, Material Constraints -34.2
Signals to Watch
- No publishable anomalous signals were retained this time.
When reading this type of Smoke briefing, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether its Integrity rating moves from pass to warn or fail. Large single-day changes in execution or constraint scores may come from question sampling, or may be early signals of real degradation, and require follow-up runs for review.
Data source: YZ Index | Run #337 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接