The 2026-09-16 YZ Index Smoke Quick Test covered 10 models, with Claude Sonnet 4.6 ranking first that day with 100 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals; it is not equivalent to a conclusion from the Full weekly ranking.
This Smoke evaluation covered only two main leaderboard dimensions: Code Execution and Material Constraints. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraints. Because the daily sample size is small, single-day scores are better used as monitoring signals rather than as long-term conclusions about model capabilities.
Daily Ranking
| Rank | Model | Main Leaderboard | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | Claude Sonnet 4.6 | 100 | 100 | 100 | pass |
| #2 | Doubao Pro | 84.55 | 71.9 | 100 | pass |
| #3 | Claude Opus 4.7 | 84 | 75 | 95 | pass |
| #4 | GPT-5.5 | 84 | 75 | 95 | pass |
| #5 | Grok 4 | 81.42 | 70.3 | 95 | pass |
| #6 | Qwen3 Max | 80.59 | 68.8 | 95 | pass |
| #7 | GPT-o3 | 70.64 | 70.5 | 70.8 | pass |
| #8 | Gemini 3.1 Pro | 70.25 | 50 | 95 | pass |
| #9 | DeepSeek V4 Pro | 67.39 | 44.8 | 95 | pass |
| #10 | Gemini 2.5 Pro | 63.42 | 47.8 | 82.5 | pass |
Data Interpretation
In today's YZ Index Smoke Quick Test, Claude Sonnet 4.6 had Code Execution 100 and Material Constraints 100, supporting a Main Leaderboard score of 100, showing a balanced high-level pairing of Code Execution and Material Constraints. Doubao Pro had Code Execution 71.9 and Material Constraints 100, with a Main Leaderboard score of 84.55, reflecting that Material Constraints significantly lifted its overall ranking. Claude Opus 4.7 and GPT-5.5 both had a Main Leaderboard score of 84, with the same structure of Code Execution 75 and Material Constraints 95, indicating that their balance between Code Execution and Material Constraints was similar. Grok 4 had Code Execution 70.3 and Material Constraints 95, with a Main Leaderboard score of 81.42; Qwen3 Max had Code Execution 68.8 and Material Constraints 95, with a Main Leaderboard score of 80.59. In the top range, Material Constraints generally remained above 95.
Gemini 2.5 Pro's Main Leaderboard score was 63.42, down 36.6 points from the previous same-scope run; its Code Execution was 47.8 and Material Constraints 82.5, down 52.2 points and 17.5 points, respectively. Claude Sonnet 4.6's Material Constraints rose 50 points from the previous run, and its Main Leaderboard score rose 22.5 points. GPT-o3 had Code Execution 70.5 and Material Constraints 70.8, with a Main Leaderboard score of 70.64, down 18.1 points from the previous run, while Code Execution fell 29.5 points. Qwen3 Max's Material Constraints rose 55.7 points from the previous run, its Main Leaderboard score rose 7.9 points, and Code Execution fell 31.2 points. Grok 4's Code Execution fell 29.7 points and Material Constraints rose 20 points, while its Main Leaderboard score fell 7.3 points.
The above score changes may stem from single-day question sampling fluctuations, or may reflect real differences in model performance on specific dimensions; they need to be confirmed by later same-scope runs. The Smoke Quick Test is a small-sample, single-day signal; the current data is for same-day observation only and does not constitute a basis for long-term trend judgments.
Key Changes
- Gemini 2.5 Pro: Main Leaderboard down 36.6 points, Code Execution -52.2 points, Material Constraints -17.5 points
- Claude Sonnet 4.6: Main Leaderboard up 22.5 points, Material Constraints +50 points
- GPT-o3: Main Leaderboard down 18.1 points, Code Execution -29.5 points
- Qwen3 Max: Main Leaderboard up 7.9 points, Code Execution -31.2 points, Material Constraints +55.7 points
- Grok 4: Main Leaderboard down 7.3 points, Code Execution -29.7 points, Material Constraints +20 points
Signals to Watch
- No publishable anomaly signals were retained for this run.
When reading this type of Smoke brief, the focus should be on two questions: first, whether a model has repeatedly exposed the same type of weakness over multiple consecutive days; second, whether its Integrity rating moves from pass to warn or fail. Large single-day changes in execution or constraint scores may come from question sampling, or may be an early signal of genuine degradation, and require follow-up runs for confirmation.
Data: YZ Index | Run #325 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接