The 2026-08-12 YZ Index Smoke quick test covered 11 models, with Grok 4 ranking first that day with 98.41 points. Smoke is a daily 10-question quick test suited for observing short-term signals, and is not equivalent to the Full weekly ranking conclusions.
This Smoke evaluation only covers two main leaderboard dimensions — Code Execution and Material Constraint — with the main leaderboard formula being 0.55 × Code Execution + 0.45 × Material Constraint. Due to the small daily sample size, single-day scores are better suited as monitoring signals rather than long-term conclusions about model capability.
Daily Rankings
| Rank | Model | Main Score | Code Execution | Material Constraint | Integrity |
|---|---|---|---|---|---|
| #1 | Grok 4 | 98.41 | 97.1 | 100 | pass |
| #2 | Gemini 2.5 Pro | 88.53 | 99.6 | 75 | pass |
| #3 | Claude Opus 4.7 | 86.03 | 74.6 | 100 | pass |
| #4 | GLM-4.6 | 79.38 | 62.5 | 100 | pass |
| #5 | Claude Sonnet 4.6 | 75 | 75 | 75 | pass |
| #6 | Gemini 3.1 Pro | 75 | 75 | 75 | pass |
| #7 | GPT-o3 | 75 | 75 | 75 | pass |
| #8 | Doubao Pro | 74.62 | 74.3 | 75 | pass |
| #9 | DeepSeek V4 Pro | 70.12 | 74.3 | 65 | pass |
| #10 | Qwen3 Max | 62.71 | 73.1 | 50 | pass |
| #11 | GPT-5.5 | 61.25 | 50 | 75 | pass |
Data Analysis
In today's YZ Index Smoke quick test, the leading models showed clear differentiation in how they combined Code Execution and Material Constraint scores. Grok 4 ranked first with a main score of 98.41, with Code Execution at 97.1 and Material Constraint at 100, maintaining high performance on both capabilities. Gemini 2.5 Pro scored 88.53 on the main leaderboard, with Code Execution at 99.6 but Material Constraint at only 75, highlighting its strength in the Code Execution dimension. Claude Opus 4.7 scored 86.03 on the main leaderboard, with Material Constraint at 100 and Code Execution at 74.6, reflecting a relative advantage in Material Constraint. GLM-4.6 scored 79.38 on the main leaderboard, with Code Execution at 62.5 and Material Constraint at 100, similarly relying on its Material Constraint score to support overall performance.
Multiple models showed notable score changes. GPT-5.5 dropped 21.9 points on the main leaderboard, with Code Execution down 50 points and Material Constraint up 12.4 points; Gemini 3.1 Pro dropped 19.7 points on the main leaderboard, with Code Execution down 25 points and Material Constraint down 13.3 points; Qwen3 Max dropped 18.9 points on the main leaderboard, with Code Execution down 14.4 points and Material Constraint down 24.3 points; DeepSeek V4 Pro dropped 17.9 points on the main leaderboard, with Code Execution down 25.7 points and Material Constraint down 8.3 points; Claude Sonnet 4.6 dropped 10.7 points on the main leaderboard, with Code Execution down 25 points and Material Constraint up 6.7 points. These single-day fluctuations may stem from differences in question sampling, or may reflect instability in specific model capabilities, and require confirmation through subsequent runs using the same methodology.
Regarding anomaly signals, GLM-4.6's integrity rating dropped to Fail, with the previous record showing fail→pass. This change appeared in a small-sample single-day test, and the specific cause still requires more run data to verify, in order to distinguish random fluctuations from genuine regression. The overall analysis is based solely on the day's score structure, avoiding conclusions about long-term model performance.
Key Changes
- GPT-5.5: Main leaderboard down 21.9 points, Code Execution -50 points, Material Constraint +12.4 points
- Gemini 3.1 Pro: Main leaderboard down 19.7 points, Code Execution -25 points, Material Constraint -13.3 points
- Qwen3 Max: Main leaderboard down 18.9 points, Code Execution -14.4 points, Material Constraint -24.3 points
- DeepSeek V4 Pro: Main leaderboard down 17.9 points, Code Execution -25.7 points, Material Constraint -8.3 points
- Claude Sonnet 4.6: Main leaderboard down 10.7 points, Code Execution -25 points, Material Constraint +6.7 points
Signals to Watch
- GLM-4.6: Integrity rating dropped to Fail (fail→pass)
When reading this type of Smoke briefing, the focus should be on two questions: first, whether a model has exposed the same type of weakness on multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large single-day changes in execution or constraint scores may come from question sampling, or may be early signals of genuine regression, requiring follow-up runs for verification.
Data source: YZ Index | Run #275 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接