From August 6 to August 9, 2026, Claude Sonnet 4.6 scored 92.17 on the first day and 69.56 on the last day, a trend of -22.6, with an average of 79.6 and volatility of 22.6.
Score Trajectories and Mechanisms of Declining Models
Claude Sonnet 4.6 posted the largest single-model decline over the four days; a first-to-last-day gap of 22.6 points directly pulled its average down to 79.6. DeepSeek V4 Pro also showed a decline, from 92.17 on the first day to 86.5 on the last, a trend of -5.7, with an average of 83.5 and volatility of 32.9. Gemini 2.5 Pro fell from 71.3 to 63.95, a trend of -7.3, with an average of 76.4 and volatility of 23.3. GLM-4.6 dropped 22.1 points, from 59.86 on the first day to 37.76 on the last, with an average of only 24.4 and volatility of 59.9; its integrity rating changed from pass to missing records.
These declines were mainly concentrated in the execution and material constraint dimensions of the main leaderboard. The consecutive low-score days for Claude Sonnet 4.6 and GLM-4.6 point to insufficient consistency in consecutive rapid testing. GLM-4.6's volatility of 59.9 indicates an extremely high daily standard deviation, with scores fluctuating widely across multiple responses to similar questions—rather than a pure accuracy problem.
Score Trajectories and Mechanisms of Rising Models
GPT-o3 rose from 89.2 to 95.91, a trend of +6.7, with an average of 88.8 and volatility of 25.3. Claude Opus 4.7 rose from 87.36 to 92.85, a trend of +5.5, with an average of 89.7 and volatility of 16.7. Gemini 3.1 Pro posted the largest gain, rising from 73.61 to 84.66, a trend of +11.1, with an average of 80.5 and volatility of 28.7. Qwen3 Max rose from 80.48 to 85.82, a trend of +5.3, with an average of 81.9 and volatility of 30.8. Grok 4 rose from 82.99 to 84.27, a trend of +1.3, with an average of 85.9 and volatility of 14.1.
The rises in GPT-o3 and Claude Opus 4.7 came with relatively low volatility, indicating that both maintained relatively stable output quality in the two main leaderboard dimensions of code execution and material constraints. Although Gemini 3.1 Pro posted the largest gain, its volatility of 28.7 remains higher than Grok 4's 14.1, suggesting that its consistency has not yet been fully consolidated during its upward trajectory.
Independent Signals from Integrity Rating Changes
GPT-o3's integrity rating changed from pass to warn, Grok 4 went from pass to warn and back to pass, and GLM-4.6's integrity rating records are missing. The integrity rating is an entry threshold rather than a bonus item; a warn status directly affects a model's deployability in production environments. Although GPT-o3 ranks near the top with an average of 88.8, its integrity rating fluctuation already constitutes a usage barrier.
Implications for Users of High-Volatility Models
Models such as GLM-4.6 (volatility 59.9), DeepSeek V4 Pro (32.9), and Qwen3 Max (30.8) are suitable for exploratory tasks with lower consistency requirements. Teams with heavy code-execution needs that adopt GLM-4.6 would need to invest in additional manual verification to cope with its drastic daily score swings. For scenarios sensitive to material fidelity, priority should go to Grok 4 (volatility 14.1) or Claude Opus 4.7 (volatility 16.7).
Strategic Judgment on Steadily Rising Models
GPT-o3 and Claude Opus 4.7 both maintained positive trends over the four days with averages above 88 points; the data supports classifying them as the stable riser group in the current Smoke rapid test. Gemini 3.1 Pro gained 11.1 points but had volatility of 28.7, and the next cycle will need to confirm whether its consistency converges. Doubao Pro recorded an average of 0 for four consecutive days, indicating it generated no valid score records in this Smoke evaluation.
Based on the current first-to-last-day comparison, the declines of Claude Sonnet 4.6 and GLM-4.6 have exceeded the gains of most rising models, and the data indicates that the probability of these two being relatively undervalued in this week's Smoke rapid test is low. GPT-o3's combination of a +6.7 trend and an average of 88.8 is the only instance this period that simultaneously satisfies both an upward trend and a relatively high average.
Data source: YZ Index | Run #270 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接