In the Smoke evaluation from 2026-08-10 to 2026-08-16, Claude Opus 4.7 ranked first among all models with a weekly average of 82.3 and a volatility of 20.7, with scores rising steadily from 72.3 on the first day to 86.25 on the last day.
Comparative Execution Performance of Rising Models
Claude Opus 4.7 posted a weekly average of 82.3, volatility of 20.7, and a trend of +14. GPT-o3 averaged 74.8, with volatility of 57.2 and a trend of +15. GPT-5.5 averaged 72.1, with volatility of 42.9 and a trend of +27.9. Claude Sonnet 4.6 averaged 75.9, with volatility of 36.5 and a trend of +19.7. Qwen3 Max averaged 69.2, with volatility of 28.2 and a trend of +10. Grok 4 averaged 79.4, with volatility of 41 and a trend of +12.7.
The combination of low volatility and a high average for Claude Opus 4.7 indicates the strongest consistency across its daily 10-question responses. Although GPT-5.5 shows the largest trend, its average of only 72.1 suggests the rise stems mainly from recovery off a low base rather than an overall capability leap.
Data Facts for Declining Models
DeepSeek V4 Pro scored 87.2 on day one and 70.44 on the last day, with a trend of -16.8, an average of 64.2, and volatility of 88. Gemini 2.5 Pro scored 72.86 on day one and 54.38 on the last day, with a trend of -18.5, an average of 62.1, and volatility of 66.3. GLM-4.6 scored 61.25 on day one and 48.75 on the last day, with a trend of -12.5, an average of 56.9, and volatility of 57.1.
DeepSeek V4 Pro's volatility of 88 is the highest among all models, indicating extreme day-to-day score variations, with sharp pullbacks likely after high-scoring days. Gemini 2.5 Pro posted the largest decline, with its final-day score falling 7.72 points below its weekly average.
Analysis of Causes for High-Volatility Models
Doubao Pro scored 0 on day one and 70.44 on the last day, with a trend of +70.4, an average of 49.2, and volatility of 78.4. DeepSeek V4 Pro posted volatility of 88. Gemini 2.5 Pro posted volatility of 66.3. GLM-4.6 posted volatility of 57.1.
High volatility typically stems from unstable execution-dimension scores across the 10 daily questions in Smoke. Both DeepSeek V4 Pro and Doubao Pro exceed 78 in volatility, meaning their code execution or material-constraint performance varies significantly from day to day, failing to sustain consistent output quality.
Independent Signals from Integrity Rating Changes
GLM-4.6 experienced a pass→fail→pass integrity rating fluctuation within 7 days. Doubao Pro had no rating records for the first two days, then shifted to pass for the remaining five days. All other models maintained a pass integrity rating throughout, with no fail occurrences.
GLM-4.6's fail day occurred during its declining trend, indicating notable problems in material constraints or task expression that directly dragged its weekly average down to 56.9. Doubao Pro's missing integrity records for the first two days may be related to incomplete participation in the evaluation during the initial phase.
Specific Implications for Users
Teams prioritizing code execution should opt for Claude Opus 4.7 first, as its 82.3 average and 20.7 volatility provide the strongest consistency guarantee. Those relying on material fidelity should avoid GLM-4.6, since its integrity rating has recorded a fail, posing a risk of unreliable output.
Developers seeking cost-effectiveness may consider Qwen3 Max, which offers better cost control among rising models with a 69.2 average and 28.2 volatility. Models with volatility above 70, such as DeepSeek V4 Pro and Doubao Pro, are unsuitable for production environments requiring stable daily output.
Strategic Assessment
Claude Opus 4.7's combination of low volatility and a high average stands out most prominently in the current data, making it suitable as a benchmark model for long-term observation. The simultaneous significant declines of DeepSeek V4 Pro and Gemini 2.5 Pro indicate persistently weakening execution dimensions across the 7-day small-sample test.
Although GPT-5.5 shows a substantial trend of 27.9, its average remains 10.2 points below Claude Opus 4.7, indicating that its upward momentum has not yet translated into a leading advantage. GLM-4.6's fluctuating integrity rating makes it the only model to clearly record a fail this week, and its material-constraint dimension needs focused verification in the next round of data to confirm whether stability has been restored.
Data source: YZ Index | Run #280 | View Raw Data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接