On 2026-09-19, the YZ Index Smoke quick test covered 10 models, and Claude Opus 4.7 ranked first for the day with 92.49 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals, and is not equivalent to Full weekly leaderboard conclusions.
This Smoke evaluation covered only two main leaderboard dimensions: code execution and material constraints. The main leaderboard formula is 0.55 × code execution + 0.45 × material constraints. Because the daily sample size is small, single-day scores are better suited as monitoring signals rather than long-term conclusions about model capability.
Daily Ranking
| Rank | Model | Main Leaderboard | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | Claude Opus 4.7 | 92.49 | 100 | 83.3 | pass |
| #2 | Grok 4 | 88.33 | 96.7 | 78.1 | warn |
| #3 | Claude Sonnet 4.6 | 78.72 | 96 | 57.6 | pass |
| #4 | GPT-o3 | 77.57 | 75 | 80.7 | pass |
| #5 | GPT-5.5 | 77.23 | 83.3 | 69.8 | pass |
| #6 | Qwen3 Max | 77.01 | 98.7 | 50.5 | warn |
| #7 | Doubao Pro | 76.88 | 96 | 53.5 | pass |
| #8 | Gemini 2.5 Pro | 72.22 | 73.7 | 70.4 | pass |
| #9 | Gemini 3.1 Pro | 58.91 | 50 | 69.8 | pass |
| #10 | DeepSeek V4 Pro | 55.02 | 25 | 91.7 | pass |
Key Changes
- GPT-o3: main leaderboard down 21.4 points, code execution -25 points, material constraints -17.1 points
- Gemini 2.5 Pro: main leaderboard down 18.7 points, code execution -26.3 points, material constraints -9.4 points
- Claude Sonnet 4.6: main leaderboard down 15.6 points, material constraints -30.3 points
- GPT-5.5: main leaderboard down 15 points, code execution -11.7 points, material constraints -19.1 points
- DeepSeek V4 Pro: main leaderboard down 12.5 points, code execution -25 points
Signals to Watch
- No publishable anomaly signals were retained in this run.
When reading this type of Smoke brief, focus on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether its integrity rating has moved from pass to warn or fail. Large single-day changes in execution or constraint scores may come from question sampling, or they may be early signals of real degradation and require subsequent runs for review.
Data source: YZ Index | Run #329 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接