The 2026-08-14 YZ Index Smoke quick test covered 11 models, with Claude Opus 4.7 ranking first for the day at 72.21 points. Smoke is a daily 10-question quick test suited for observing short-term signals and is not equivalent to the Full weekly leaderboard conclusions.
This Smoke evaluation covers only two main leaderboard dimensions: code execution and material constraint. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraint. Given the small daily sample size, single-day scores are better used as monitoring signals rather than long-term conclusions about model capability.
Daily Rankings
| Rank | Model | Main Leaderboard | Code Execution | Material Constraint | Integrity |
|---|---|---|---|---|---|
| #1 | Claude Opus 4.7 | 72.21 | 70 | 74.9 | pass |
| #2 | Claude Sonnet 4.6 | 68.11 | 70 | 65.8 | pass |
| #3 | Grok 4 | 57.42 | 45 | 72.6 | pass |
| #4 | GPT-5.5 | 54.36 | 45 | 65.8 | pass |
| #5 | Qwen3 Max | 53.33 | 45 | 63.5 | pass |
| #6 | Doubao Pro | 52.3 | 61.7 | 40.8 | pass |
| #7 | DeepSeek V4 Pro | 49.8 | 36.7 | 65.8 | pass |
| #8 | GPT-o3 | 40.1 | 45 | 34.1 | pass |
| #9 | Gemini 3.1 Pro | 38.55 | 36.7 | 40.8 | pass |
| #10 | Gemini 2.5 Pro | 22.25 | 20 | 25 | pass |
| #11 | GLM-4.6 | 22.25 | 20 | 25 | pass |
Data Interpretation
In today's Smoke quick test, the top models showed differences in their combination of code execution and material constraint. Claude Opus 4.7 scored 72.21 on the main leaderboard, with code execution at 70 and material constraint at 74.9; Claude Sonnet 4.6 scored 68.11, with code execution at 70 and material constraint at 65.8. Both maintained relatively high code execution levels. Grok 4 scored 57.42 on the main leaderboard, with code execution at 45 and material constraint at 72.6, showing relatively strong material constraint; GPT-5.5 scored 54.36, also with code execution at 45 and material constraint at 65.8, leaning toward material constraint in structure.
Multiple models saw significant score declines: GPT-o3 dropped 57.2 points on the main leaderboard, -50 in code execution, and -65.9 in material constraint; Gemini 2.5 Pro dropped 57.2 points on the main leaderboard, -50 in code execution, and -65.9 in material constraint; GLM-4.6 dropped 56.2 points on the main leaderboard, -50 in code execution, and -63.7 in material constraint; Gemini 3.1 Pro dropped 45 points on the main leaderboard, -33.3 in code execution, and -59.2 in material constraint; GPT-5.5 dropped 42.9 points on the main leaderboard, -50 in code execution, and -34.2 in material constraint. These changes could stem from question sampling fluctuation or could be genuine degradation signals, requiring confirmation in subsequent runs.
As a small-sample single-day signal, Smoke only reflects the immediate combination of code execution and material constraint for the day. We keep our wording measured and avoid drawing long-term conclusions.
Key Changes
- GPT-o3: Main leaderboard down 57.2 points, code execution -50, material constraint -65.9
- Gemini 2.5 Pro: Main leaderboard down 57.2 points, code execution -50, material constraint -65.9
- GLM-4.6: Main leaderboard down 56.2 points, code execution -50, material constraint -63.7
- Gemini 3.1 Pro: Main leaderboard down 45 points, code execution -33.3, material constraint -59.2
- GPT-5.5: Main leaderboard down 42.9 points, code execution -50, material constraint -34.2
Signals to Watch
- No publishable anomaly signals were retained this time.
When reading Smoke briefings like this, the focus should be on two questions: first, whether a model has exposed the same type of weakness across multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large single-day swings in execution or constraint scores may come from question sampling or may be early signs of genuine degradation, requiring review in subsequent runs.
Data source: YZ Index | Run #278 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接