The 2026-10-12 YZ Index Smoke quick test covered 14 models, with GPT-6 Astra and GPT-6.1 Sol tied for first place for the day at 84.99. Smoke is a daily 10-question quick test, well suited for observing short-term signals, and is not equivalent to Full weekly leaderboard conclusions.
This Smoke evaluation covers only two main leaderboard dimensions: Code Execution and Material Constraints. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraints. Because the daily sample size is small, single-day scores are better used as monitoring signals rather than as long-term conclusions about model capability.
Daily Ranking
| Ranking | Model | Main | Code Execution | Material Constraints | Integrity |
|---|---|---|---|---|---|
| #1 | GPT-6 Astra | 84.99 | 97 | 70.3 | pass |
| #2 | GPT-6.1 Sol | 84.99 | 97 | 70.3 | pass |
| #3 | Doubao Pro | 84.66 | 91.5 | 76.3 | pass |
| #4 | GPT-6 Sol | 82.79 | 94.8 | 68.1 | pass |
| #5 | GPT-6 Luna | 78.55 | 97 | 56 | pass |
| #6 | Claude Sonnet 4.6 | 73.94 | 72 | 76.3 | pass |
| #7 | DeepSeek V4 Pro | 72.59 | 72 | 73.3 | pass |
| #8 | GPT-5.5 | 72.59 | 72 | 73.3 | pass |
| #9 | Gemini 3.1 Pro | 71.6 | 72 | 71.1 | pass |
| #10 | GPT-o3 | 71.38 | 69.8 | 73.3 | pass |
| #11 | Gemini 2.5 Pro | 71.21 | 69.5 | 73.3 | pass |
| #12 | Grok 4 | 68.02 | 63.7 | 73.3 | pass |
| #13 | Claude Opus 4.7 | 58.84 | 47 | 73.3 | pass |
| #14 | GLM-4.6 | 41.18 | 16.7 | 71.1 | pass |
Data Interpretation
Today's top two on the main leaderboard, GPT-6 Astra and GPT-6.1 Sol, both show a structure of Code Execution 97 and Material Constraints 70.3, with both main leaderboard scores at 84.99; Doubao Pro has Code Execution 91.5 and Material Constraints 76.3 for a main leaderboard score of 84.66, showing a combination with relatively stronger Material Constraints. GPT-6 Luna has Code Execution 97 and Material Constraints 56 for a main leaderboard score of 78.55, reflecting a Code Execution-dominated profile. Claude Sonnet 4.6 has Code Execution 72 and Material Constraints 76.3 for a main leaderboard score of 73.94, contrasting with the Material Constraints score of 73.3 in the same band as DeepSeek V4 Pro and GPT-5.5.
Doubao Pro gained 29.9 points on the main leaderboard, with Code Execution +44.5 and Material Constraints +12; GPT-5.5 gained 24.2 points on the main leaderboard, with Code Execution +25 and Material Constraints +23.3; DeepSeek V4 Pro gained 23.1 points on the main leaderboard, with Code Execution +25 and Material Constraints +20.7; GPT-6 Sol gained 16.8 points on the main leaderboard, with Code Execution +47.8 and Material Constraints -21.2, reflecting pronounced daily fluctuations in the pairing of Code Execution and Material Constraints. GPT-6 Astra's Material Constraints plunged -19 points, GPT-6.1 Sol's Material Constraints plunged -19 points, Gemini 2.5 Pro's Material Constraints plunged -16 points, Claude Opus 4.7's main leaderboard score plunged -21 points, and GLM-4.6's main leaderboard score plunged -13.6 points; these anomalous signals may stem from question sampling fluctuation or may be signs of genuine degradation, requiring subsequent runs for verification.
Smoke is a small-sample single-day signal; the above interpretation is based only on that day's data structure and does not constitute a long-term judgment.
Key Changes
- Doubao Pro: Main leaderboard up 29.9 points, Code Execution +44.5, Material Constraints +12
- GPT-5.5: Main leaderboard up 24.2 points, Code Execution +25, Material Constraints +23.3
- DeepSeek V4 Pro: Main leaderboard up 23.1 points, Code Execution +25, Material Constraints +20.7
- Claude Opus 4.7: Main leaderboard down 21 points, Code Execution -25, Material Constraints -16
- GPT-6 Sol: Main leaderboard up 16.8 points, Code Execution +47.8, Material Constraints -21.2
Signals to Watch
- GPT-6 Astra: Material Constraints plunged -19 points
- GPT-6.1 Sol: Material Constraints plunged -19 points
- GPT-6 Sol: Material Constraints plunged -21.2 points
- Gemini 2.5 Pro: Material Constraints plunged -16 points
- Claude Opus 4.7: Main leaderboard plunged -21 points
- GLM-4.6: Main leaderboard plunged -13.6 points
- Qwen3 Max: Incomplete data (missing execution, evidence, judgment, integrity, and communication dimensions; API failure/timeout), has entered automatic rerun, and is not included in this period's ranking
When reading this type of Smoke brief, focus on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether its integrity rating has moved from pass to warn or fail. Large changes in single-day execution or constraint scores may come from question sampling or may be early signals of genuine degradation, requiring follow-up runs to verify.
Data source: YZ Index | Run #373 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接