During 2026-09-07 to 2026-09-13, Grok 4 rose from 80.13 points on the first day to 87 points on the last day, a trend of +6.9, an average of 85.8 points, and volatility of 17.2; GLM-4.6 fell from 83.49 points on the first day to 0 points on the last day, a trend of -83.5, an average of 20.7 points, and volatility of 83.5.
Data and Mechanisms of Rising Models
Grok 4 scored 80.13 on the first day and 87 on the last day, with a trend of +6.9, an average of 85.8, and volatility of 17.2. GPT-o3 scored 77.5 on the first day and 85.15 on the last day, with a trend of +7.7, an average of 80.4, and volatility of 32.2. Doubao Pro scored 80.85 on the first day and 85.15 on the last day, with a trend of +4.3, an average of 86.3, and volatility of 20.9. Claude Opus 4.7 scored 61.25 on the first day and 63.59 on the last day, with a trend of +2.3, an average of 71.4, and volatility of 23.5. Qwen3 Max scored 55.99 on the first day and 57.65 on the last day, with a trend of +1.7, an average of 66.5, and volatility of 30.5.
Among rising models, Grok 4 and GPT-o3 had trend gains above 6 points, while Doubao Pro had the highest average. Grok 4 and Doubao Pro, with volatility below 21, showed relatively strong answer consistency. GPT-o3 had volatility of 32.2, higher than Grok 4, indicating larger day-to-day score swings.
Data and Mechanisms of Declining Models
DeepSeek V4 Pro scored 83.49 on the first day and 77.34 on the last day, with a trend of -6.1, an average of 77.9, and volatility of 23. Gemini 3.1 Pro scored 83.49 on the first day and 73.25 on the last day, with a trend of -10.2, an average of 77.3, and volatility of 39.4. GLM-4.6 scored 83.49 on the first day and 0 on the last day, with a trend of -83.5, an average of 20.7, and volatility of 83.5. GPT-5.5 scored 72.03 on the first day and 60.87 on the last day, with a trend of -11.2, an average of 75.9, and volatility of 34.4. Gemini 2.5 Pro scored 71.16 on the first day and 59.5 on the last day, with a trend of -11.7, an average of 74.4, and volatility of 29.3. Claude Sonnet 4.6 scored 65.24 on the first day and 53.56 on the last day, with a trend of -11.7, an average of 73.1, and volatility of 32.7.
Among declining models, GLM-4.6 had volatility of 83.5, far higher than the others, while Gemini 3.1 Pro ranked second with volatility of 39.4. GLM-4.6's last-day score dropped directly to zero; combined with its integrity rating starting at warn and subsequently blank, this indicates extremely low answer consistency. Gemini 3.1 Pro and GPT-5.5 both had volatility above 34 and trend declines above 10 points.
Separate Analysis of Integrity Ratings
GLM-4.6's integrity rating was warn → warn → blank, with consecutive non-pass states, directly corresponding to its volatility of 83.5. Grok 4's integrity rating was pass → warn → warn → pass → pass → pass → pass; after two days of warn, it returned to pass, and its volatility was only 17.2, lower than most declining models. Periods of warn integrity ratings correlate with widening score volatility, and GLM-4.6's blank records further confirm its unstable answers.
Scenario Implications for Users
For teams with code-heavy execution scenarios, Grok 4, with an average of 85.8 and volatility of 17.2, is suitable as a primary model; GPT-o3, with an average of 80.4 but volatility of 32.2, requires backup options to handle day-to-day fluctuations. In scenarios sensitive to material constraints, Doubao Pro has the highest average at 86.3 and can be prioritized. GLM-4.6 has an average of only 20.7 and a last-day score of 0, posing a direct risk to developers who rely on stable output.
The engineering-judgment side of the leaderboard shows that models with volatility above 30 have significantly different scores across the 7-day test; Gemini 3.1 Pro had volatility of 39.4 and GPT-5.5 had volatility of 34.4, both higher than the average level of rising models. In the task-expression side of the leaderboard, models with a warn integrity rating require additional manual review of outputs.
Strategic Assessment
Based on 2026-W37 data, the combination of low volatility and high averages for Grok 4 and Doubao Pro stands out in the current sample; GLM-4.6, with high volatility and integrity issues, has the highest risk of being underestimated. Although Gemini 3.1 Pro tied DeepSeek V4 Pro at 83.49 on the first day, it was 3.91 points lower on the last day, indicating insufficient consistency. The next period needs to verify whether Grok 4 maintains a trend gain above +6.9 and whether GLM-4.6 continues blank records.
Claude Opus 4.7 had a trend of +2.3 but an average of 71.4, making it suitable for scenarios with moderate stability requirements; Qwen3 Max had a trend of +1.7, an average of 66.5, and volatility of 30.5, so it remains to be seen whether it can narrow the 20-point gap with leading models. All judgments come directly from first-day and last-day scores, trend values, averages, and volatility data.
Data source: YZ Index | Run #320 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接