During the period from 2026-09-14 to 2026-09-20, Claude Opus 4.7 rose from 91.09 points on the first day to 96.99 points on the last day, with a trend value of +5.9 and an average of 87.9 points.
Score Trajectories and Mechanisms of Rising Models
Qwen3 Max had a weekly trend of +19.1, rising from 70.47 points to 89.52 points, with an average of 75.3 points and volatility of 34.9. The model showed a clear climb over seven consecutive days of testing, and its last-day score tied Claude Sonnet 4.6 at 89.52 points. Claude Sonnet 4.6 had a trend of +7, rising from 82.5 points to 89.52 points, with an average of 85.6 points and volatility of 23.6. Doubao Pro had a trend of +3.3, rising from 91.09 points to 94.35 points, with an average of 83.5 points and volatility of 24.2.
These increases were mainly reflected in the two main leaderboard dimensions: code execution and material constraints. After a low score on the first day, Qwen3 Max gradually stabilized, and its last-day score was already close to the level of Claude Opus 4.7, indicating that its response consistency on execution-type tasks in the Smoke 10-question quick test had improved.
Score Trajectories and Mechanisms of Declining Models
DeepSeek V4 Pro had a trend of -15.3, falling from 91.09 points to 75.77 points, with an average of 66.6 points and volatility of 43.3, the highest volatility value this week. Gemini 3.1 Pro had a trend of -11.6, falling from 87 points to 75.38 points, with an average of 71.2 points and volatility of 28.1. Grok 4 had a trend of -11.2, falling from 86.61 points to 75.38 points, with an average of 81 points and volatility of 23.3. GPT-5.5 had a trend of -4.6, falling from 91.09 points to 86.5 points, with an average of 84.3 points and volatility of 17.1. Gemini 2.5 Pro had a trend of -3, falling from 70.19 points to 67.24 points, with an average of 77.3 points and volatility of 36.6.
The volatility values of DeepSeek V4 Pro and Gemini 2.5 Pro both exceeded 36, indicating significant differences in their answers to the same types of questions over seven days. GLM-4.6 scored 0 throughout, with a trend of 0, an average of 0, and volatility of 0, and did not enter the valid scoring range.
Separate Analysis of Integrity Rating Changes
Doubao Pro's integrity rating showed warn on the fifth day and pass on the other six days. Grok 4 showed warn on the sixth day and pass on the others. Qwen3 Max had warn on the first day and warn on the sixth day, and pass on the others. None of the three models' warns occurred consecutively, but the warn periods for Qwen3 Max and Doubao Pro corresponded exactly to days when their score fluctuations were relatively large, indicating an association between integrity rating warn and decreased answer consistency.
Analysis of Causes for High-Volatility Models
DeepSeek V4 Pro's volatility was 43.3, Qwen3 Max's 34.9, Claude Opus 4.7's 33.2, and Gemini 2.5 Pro's 36.6, all higher than those of other models. Smoke has only 10 questions per day; although the sample is small, seven consecutive days of data already show that these models have relatively large score standard deviations in the code execution and material constraint dimensions. Under the stability formula max(0,100-stddev×2), high volatility directly lowers the stability score, reflecting insufficient answer consistency.
Specific Implications for Users
For teams that focus heavily on code execution, if they rely on DeepSeek V4 Pro, its average of 66.6 points and volatility of 43.3 mean that daily task result differences may exceed 30 points, requiring additional manual review steps. For scenarios sensitive to material fidelity, the combination of Claude Opus 4.7's last-day score of 96.99 points and average of 87.9 points is more useful as a reference, but its volatility of 33.2 still requires developers to perform multiple rounds of validation before critical outputs.
For Doubao Pro and Qwen3 Max, whose integrity ratings showed warn, developers should add extra fact-checking steps in production environments to prevent low-consistency outputs during warn periods from entering downstream systems.
Strategic Assessment
Based on this week's data, Qwen3 Max and Claude Sonnet 4.6 had trend values of +19.1 and +7, respectively, and their last-day scores have entered a high range, suggesting relative to their averages that they may be undervalued. DeepSeek V4 Pro and Gemini 3.1 Pro had negative trend values exceeding -11 and were among the highest in volatility, suggesting their current scores may be overvalued. GLM-4.6's record of 0 points indicates that it did not produce valid output in this round of Smoke testing, and next period's verification is needed to determine whether it recovers.
The integrity rating warn signal appeared in three models, suggesting that enterprises selecting models should use integrity rating as an admission threshold and prioritize excluding models with higher warn frequency. All judgments come directly from the 7-day scores, trends, and volatility data for 2026-W38 and do not exceed the scope supported by the data.
Data source: YZ Index | Run #330 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接