The 2026-09-15 YZ Index Smoke quick test covered 10 models, with Gemini 2.5 Pro ranking first for the day at 100 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals, and is not equivalent to the conclusions of the Full weekly leaderboard.
This Smoke evaluation covered only two main leaderboard dimensions: code execution and material constraints. The main leaderboard formula is 0.55 × code execution + 0.45 × material constraints. Because the daily sample size is small, single-day scores are better used as monitoring signals rather than long-term conclusions about model capability.
Daily Ranking
| Rank | Model | Main Leaderboard | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | Gemini 2.5 Pro | 100 | 100 | 100 | pass |
| #2 | GPT-o3 | 88.75 | 100 | 75 | pass |
| #3 | Grok 4 | 88.75 | 100 | 75 | pass |
| #4 | Claude Opus 4.7 | 83.94 | 100 | 64.3 | pass |
| #5 | Doubao Pro | 83.94 | 100 | 64.3 | pass |
| #6 | GPT-5.5 | 83.94 | 100 | 64.3 | pass |
| #7 | Claude Sonnet 4.6 | 77.5 | 100 | 50 | pass |
| #8 | Qwen3 Max | 72.69 | 100 | 39.3 | pass |
| #9 | Gemini 3.1 Pro | 70.19 | 75 | 64.3 | pass |
| #10 | DeepSeek V4 Pro | 61.25 | 50 | 75 | pass |
Key Changes
- Gemini 2.5 Pro: main leaderboard up 29.8 points, code execution +25 points, material constraints +35.7 points
- DeepSeek V4 Pro: main leaderboard down 29.8 points, code execution -50 points, material constraints -5.2 points
- GPT-o3: main leaderboard up 18.6 points, code execution +25 points, material constraints +10.7 points
- Gemini 3.1 Pro: main leaderboard down 16.8 points, code execution -25 points, material constraints -6.8 points
- Claude Opus 4.7: main leaderboard down 7.2 points, material constraints -15.9 points
Signals to Watch
- Claude Opus 4.7: material constraints plunged by -15.9 points
- Doubao Pro: material constraints plunged by -15.9 points
- GPT-5.5: material constraints plunged by -15.9 points
- Qwen3 Max: material constraints plunged by -40.9 points
- Gemini 3.1 Pro: main leaderboard plunged by -16.8 points
- DeepSeek V4 Pro: main leaderboard plunged by -29.8 points
- GLM-4.6: incomplete data, with several evaluation dimensions missing due to API failure/timeout; it has entered automatic rerun and is not included in this ranking
When reading this type of Smoke brief, the focus should be on two questions: first, whether a given model exposes the same type of weakness over several consecutive days; second, whether its integrity rating moves from pass to warn or fail. Large single-day changes in execution or constraint scores may come from question sampling, or they may be early signals of real degradation, and require follow-up runs for review.
Data source: YZ Index | Run #324 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接