Qwen3 Max Rallies +36.8 to Lead; Gemini 3.1 Pro Slips 5.6 as Biggest Loser

Between July 28 and August 2, 2026, Qwen3 Max rose from 59.28 on the first day to 96.1 on the last day in the Smoke evaluation, with a trend value of +36.8 — the largest increase among all models over the seven-day period.

Scoring Trajectories and Mechanisms of Rising Models

Doubao Pro rose from 92.07 to 96.7, with a trend of +4.6, a mean of 84.9, and volatility of 21.7. Grok 4 rose from 84.47 to 96.1, with a trend of +11.6, a mean of 82.5, and volatility of 26.5. DeepSeek V4 Pro rose from 84 to 94.45, with a trend of +10.5, a mean of 87.8, and volatility of 20.1. Claude Sonnet 4.6 rose from 79.91 to 94.45, with a trend of +14.5, a mean of 75.5, and volatility of 39.4. Gemini 2.5 Pro rose from 77.6 to 96.1, with a trend of +18.5, a mean of 70.6, and volatility of 37.2. Claude Opus 4.7 rose from 56.5 to 94.45, with a trend of +38, a mean of 83.4, and volatility of 38.

These gains were mainly concentrated in the two main-leaderboard dimensions of execution and material constraints. DeepSeek V4 Pro's mean of 87.8 was the highest among rising models, indicating relatively better day-to-day consistency in code execution and material constraints. Qwen3 Max and Claude Opus 4.7 reached final-day scores of 96.1 and 94.45, respectively, showing that they gradually narrowed the gap with top models over seven consecutive days.

Scoring Trajectories and Mechanisms of Declining Models

Gemini 3.1 Pro dropped from a first-day 100 to a final-day 94.45, with a trend of -5.6, a mean of 83.9, and volatility of 27.3. GPT-o3 dropped from 97.75 to 96.1, with a trend of -1.7, a mean of 87.6, and volatility of 18.5. GLM-4.6 dropped from 77.6 to 74, with a trend of -3.6, a mean of 59.7, and volatility of 38.

Gemini 3.1 Pro continued to decline after a perfect first-day score, indicating multiple point deductions in the material-constraint dimension. GPT-o3's mean of 87.6 remained higher than that of most rising models, but its negative trend showed that execution-dimension stability was insufficient to offset a slight decline. GLM-4.6's mean of only 59.7 was the lowest among all models; combined with its integrity rating repeatedly showing 'warn,' this indicates that both the material-constraint and integrity dimensions simultaneously dragged down its overall score.

Analysis of Causes for High-Volatility Models

Claude Sonnet 4.6 had a volatility of 39.4, Gemini 2.5 Pro 37.2, Claude Opus 4.7 38, Qwen3 Max 36.8, and GLM-4.6 38. All of these models had stability scores below 35, meaning that when similar questions were answered multiple times, the standard deviation of scores was large.

High volatility most likely stems from inconsistent performance across different tasks in the execution dimension. The Claude-series models scored low on the first day and rebounded sharply later, indicating that their responses to material constraints varied significantly by date. Qwen3 Max and Gemini 2.5 Pro also exhibited a pattern of low early scores and high late scores, with volatility values close to 37, reflecting strong randomness in the code-execution stage.

Standalone Signals from Integrity-Rating Changes

DeepSeek V4 Pro's integrity rating fluctuated through pass→warn→pass→pass→pass→warn. GLM-4.6 likewise recorded pass→warn→pass→pass→warn. The integrity rating serves as an entry threshold, and a 'warn' status directly affects a model's availability on the engineering-judgment and task-expression side leaderboards.

DeepSeek V4 Pro's mean of 87.8 ranked among the top of all models, yet its integrity rating repeatedly showed 'warn,' indicating occasional inconsistent behavior in the material-constraint stage. GLM-4.6, with a mean of 59.7 and multiple 'warn' integrity ratings, shows systematic issues in both dimensions of the main leaderboard.

Practical Implications for Users

Teams that emphasize code execution should give priority to DeepSeek V4 Pro, whose mean of 87.8 and trend of +10.5 indicate relatively stable performance in the execution dimension. For scenarios sensitive to material fidelity, GLM-4.6 should be avoided, as its mean of 59.7 and multiple 'warn' integrity ratings carry higher risk.

Developers who rely on stable output across multiple consecutive days should focus on GPT-o3, whose volatility of 18.5 is the lowest among all models, making it suitable for production environments with high consistency requirements. Claude Opus 4.7 and Qwen3 Max have final-day scores close to 96, but with volatility exceeding 36, they are suitable for trial use in non-critical tasks.

Strategic Assessment

Gemini 3.1 Pro's mean of 83.9 remains higher than that of most rising models, but its trend of -5.6 shows that its leading position is being gradually caught up. Qwen3 Max and Claude Opus 4.7 posted trend values of +36.8 and +38, respectively, indicating that low-baseline models have significant room to catch up in both the execution and material-constraint dimensions.

GLM-4.6, with a mean of 59.7 and multiple 'warn' integrity ratings, is a model with high, easily underestimated risk. GPT-5.5 maintained a flat trend of 0.8, with a mean of 82.3 and volatility of 28.5, placing it in the middle and making it suitable as a benchmark reference.

This week's data supports only the judgments above, which are based on the seven-day scores from July 28 to August 2, 2026. The next edition will need to verify whether DeepSeek V4 Pro's integrity rating remains consistently stable and whether Qwen3 Max and Claude Opus 4.7 sustain their strong trends.


Data source: YZ Index | Run #257 | View Raw Data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!