According to Smoke evaluation data from 2026-09-21 to 2026-09-27, Claude Opus 4.7 rose from 70.77 points to 96.01 points over seven days, with a trend of 25.2 and an average of 87.8, the highest among all models.
Analysis of Steadily Rising Models
Claude Opus 4.7 scored 70.77 on the first day and 96.01 on the last, with a trend of 25.2, an average of 87.8, and volatility of 26. The model's score continued to rise over the seven days, and by the final day it tied GPT-o3 at 96.01. GPT-o3 scored 81.44 on the first day and 96.01 on the last, with a trend of 14.6 and an average of 84.3. Qwen3 Max scored 72.25 on the first day and 86.44 on the last, with a trend of 14.2 and an average of 76.3. GPT-5.5 scored 76.44 on the first day and 78.8 on the last, with a trend of 2.4 and an average of 83.1. Grok 4 scored 95.44 on the first day and 97.66 on the last, with a trend of 2.2 and an average of 82.7.
These upward trends most likely stem from steady improvement in the two main leaderboard dimensions: code execution and material constraints. Claude Opus 4.7 had the highest average, indicating that in the continuous 10-question quick test, its execution accuracy and consistency with source materials both remained at a high level.
For enterprises currently selecting models, Claude Opus 4.7 is suited to scenarios that emphasize code execution and require high material fidelity. Developers relying on this model may prioritize it for workflows that need stable output over consecutive days.
Strategic judgment: Claude Opus 4.7's average of 87.8 is notably higher than other rising models, and current data support the conclusion that it is undervalued.
Analysis of Clearly Declining Models
Gemini 3.1 Pro scored 95.19 on the first day and 68.26 on the last, with a trend of -26.9, an average of 80.2, and volatility of 26.9. DeepSeek V4 Pro scored 95.01 on the first day and 69.31 on the last, with a trend of -25.7, an average of 76.2, and volatility of 29. Gemini 2.5 Pro scored 85.63 on the first day and 69.91 on the last, with a trend of -15.7, an average of 78.5, and volatility of 27.4. Doubao Pro scored 90.62 on the first day and 83 on the last, with a trend of -7.6, an average of 84.3, and volatility of 19.4.
The declines were most likely caused by lower scores in the source-fidelity dimension. Gemini 3.1 Pro was still at a high level on the first day, but fell to 68.26 by the final day, declining continuously over the seven days, indicating a systemic pullback in material-constraint capability.
For scenarios sensitive to material fidelity, teams making heavy use of Gemini 3.1 Pro need to reassess risk. Developers relying on this model should prepare a switching plan to avoid output-quality fluctuations affecting business.
Strategic judgment: Gemini 3.1 Pro and DeepSeek V4 Pro both had trends below -25, and current data support the conclusion that both are overvalued.
Key Focus on High-Volatility Models
GLM-4.6 scored 0 on the first day and 57.01 on the last, with a trend of 57, an average of 31.1, and volatility of 73.9. Claude Sonnet 4.6 scored 72.5 on the first day and 92.19 on the last, with a trend of 19.7, an average of 79.8, and volatility of 45.1. Grok 4 had volatility of 36.4, and Qwen3 Max had volatility of 31.1.
High volatility directly reflects lower model response consistency. GLM-4.6's standard deviation led to a stability score of only 31.7, with huge score differences when answering similar questions multiple times. Claude Sonnet 4.6 had volatility of 45.1, also showing clear fluctuations in execution and source-alignment dimension scores.
For enterprises needing stable output, teams focused on code execution should avoid using GLM-4.6 in production environments. Developers relying on this model need to add manual review steps to cope with sharp changes in single-day scores.
Strategic judgment: GLM-4.6 and Claude Sonnet 4.6 had volatility far exceeding other models, and current data support the conclusion that their consistency is insufficient; the next period needs to verify whether they can reduce standard deviation.
Signals from Integrity Rating Changes
Grok 4's integrity rating remained stable after changing from warn to pass. GLM-4.6's integrity rating recovered to pass after appearing as warn, with blank records in between.
The integrity rating is an access threshold; Grok 4's change from warn to pass indicates improvement in its response integrity. GLM-4.6's rating fluctuated, suggesting possible early risk of nonstandard material citation.
For enterprises selecting models, prioritizing models with stable integrity ratings can reduce compliance risk. Developers relying on GLM-4.6 need to additionally cross-check output sources.
Strategic judgment: Grok 4's integrity rating improvement synchronized with its score rise, and current data support the conclusion that its overall reliability has improved.
Across the seven days of data, Claude Opus 4.7 had the best average and trend, Gemini 3.1 Pro and DeepSeek V4 Pro had the largest declines, GLM-4.6 and Claude Sonnet 4.6 were the most volatile, and integrity rating changes appeared only in Grok 4 and GLM-4.6.
Data source: YZ Index | Run #340 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接