On 2026-09-23, the YZ Index Smoke quick test covered 10 models, with Claude Sonnet 4.6 ranking first for the day at 96.98 points. Smoke is a daily 10-question quick test, suited to observing short-term signals; it is not equivalent to the conclusions of the Full weekly leaderboard.
This Smoke evaluation covered only two main leaderboard dimensions: code execution and material constraints. The main leaderboard formula is 0.55 × code execution + 0.45 × material constraints. Because the daily sample size is small, single-day scores are better used as monitoring signals than as long-term conclusions about model capability.
Daily Ranking
| Ranking | Model | Main Leaderboard | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | Claude Sonnet 4.6 | 96.98 | 94.5 | 100 | pass |
| #2 | Claude Opus 4.7 | 96.76 | 94.1 | 100 | pass |
| #3 | Gemini 2.5 Pro | 95.19 | 93.7 | 97 | pass |
| #4 | GPT-5.5 | 91.89 | 89.5 | 94.8 | pass |
| #5 | Qwen3 Max | 91.26 | 85.9 | 97.8 | pass |
| #6 | Grok 4 | 80.89 | 69.5 | 94.8 | pass |
| #7 | GPT-o3 | 80.48 | 64.5 | 100 | pass |
| #8 | Gemini 3.1 Pro | 79.13 | 64.5 | 97 | pass |
| #9 | DeepSeek V4 Pro | 78.14 | 64.5 | 94.8 | pass |
| #10 | Doubao Pro | 73.5 | 64.5 | 84.5 | pass |
Data Interpretation
In today's YZ Index Smoke quick test, leading models showed a balanced high-end pattern in combining code execution and material constraints. Claude Sonnet 4.6 ranked first on the main leaderboard at 96.98, with code execution 94.5 and material constraints 100; Claude Opus 4.7 had main leaderboard 96.76, code execution 94.1, material constraints 100; Gemini 2.5 Pro had main leaderboard 95.19, code execution 93.7, material constraints 97. All three maintained Integrity pass, and the material constraints dimension generally approached or reached 100, while code execution was in the 93.7 to 94.5 range, forming a structurally complementary advantage of strengths.
Several models showed notable changes under the same measurement basis. Gemini 2.5 Pro rose 22.7 points on the main leaderboard, with code execution +21.7 points and material constraints +23.9 points; Qwen3 Max rose 12.8 points on the main leaderboard, with code execution +13.9 points and material constraints +11.5 points; DeepSeek V4 Pro rose 12.1 points on the main leaderboard, with code execution -7.5 points and material constraints +36.1 points. Doubao Pro fell 17.3 points on the main leaderboard, with code execution -31.8 points; GPT-o3 fell 15.5 points on the main leaderboard, with code execution -32.5 points and material constraints +5.2 points. These changes stem from small-sample sampling of 10 questions on a single day and require follow-up runs to distinguish fluctuation from genuine degradation.
Anomalous signals were concentrated in several models. Grok 4 plunged -10 points on the main leaderboard, GPT-o3 plunged -15.5 points on the main leaderboard, Gemini 3.1 Pro plunged -31.8 points in code execution, and Doubao Pro plunged -17.3 points on the main leaderboard. GLM-4.6 had incomplete data due to an API failure and did not participate in the ranking. These are all single-day Smoke signals; the magnitude of change may be affected by question sampling, and multiple rounds of retesting are needed to confirm stability.
Major Changes
- Gemini 2.5 Pro: main leaderboard up 22.7 points, code execution +21.7 points, material constraints +23.9 points
- Doubao Pro: main leaderboard down 17.3 points, code execution -31.8 points
- GPT-o3: main leaderboard down 15.5 points, code execution -32.5 points, material constraints +5.2 points
- Qwen3 Max: main leaderboard up 12.8 points, code execution +13.9 points, material constraints +11.5 points
- DeepSeek V4 Pro: main leaderboard up 12.1 points, code execution -7.5 points, material constraints +36.1 points
Signals to Watch
- Grok 4: main leaderboard plunged -10 points
- GPT-o3: main leaderboard plunged -15.5 points
- Gemini 3.1 Pro: code execution plunged -31.8 points
- Doubao Pro: main leaderboard plunged -17.3 points
- GLM-4.6: incomplete data (missing several evaluation dimensions due to API failure/timeout), has entered automatic rerun, and is not ranked in this period
When reading this kind of Smoke brief, the focus should be on two questions: first, whether a model has repeatedly shown the same type of weakness over multiple consecutive days; second, whether the integrity rating moves from pass to warn or fail. Large single-day changes in execution or constraint scores may come from question sampling, or they may be early signals of genuine degradation, requiring follow-up runs to confirm.
Data source: YZ Index | Run #335 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接