Claude Opus 4.7 Trends Up 10.1 Points, Gemini 3.1 Pro Falls 11.2 Points: 2026-W35 Smoke Weekly Trends

In the Smoke evaluation from 2026-08-24 to 2026-08-30, Claude Opus 4.7 rose from 82.98 points on the first day to 93.08 points on the final day, a trend increase of 10.1 points.

Rising Models: Data and Mechanics

Claude Opus 4.7 trended +10.1 this week, closing at 93.08 points, with a mean of 86.5 points and volatility of 27.5 points. Claude Sonnet 4.6 trended +12.1, rising from 58.23 to 70.33 points, with a mean of 79 points and volatility of 32.6 points. DeepSeek V4 Pro trended +7.6, rising from 80.46 to 88.08 points, with a mean of 70.7 points and volatility of 39.1 points. Qwen3 Max trended +8.3, rising from 71.73 to 80.07 points, with a mean of 80.6 points and volatility of 21.8 points. Grok 4 trended +1.1, rising from 87.09 to 88.14 points, with a mean of 90 points and volatility of 13.6 points.

These gains are primarily associated with consecutive days of improved scores on the code execution dimension. Claude Opus 4.7 remained stable on the material constraint dimension while its execution dimension gradually narrowed deviations from reference answers, driving the final-day score upward. Grok 4's volatility of just 13.6 points indicates strong output consistency across the two primary leaderboard dimensions of execution and constraint.

Falling Models: Data and Mechanics

Gemini 3.1 Pro trended -11.2, falling from 96.98 points on the first day to 85.83 points on the final day, with a mean of 83.7 points and volatility of 33.4 points. Gemini 2.5 Pro trended -6.4, dropping from 92.16 to 85.73 points, with a mean of 88.2 points and volatility of 21.1 points. GPT-o3 trended -13.7, falling from 83.23 to 69.56 points, with a mean of 87.1 points and volatility of 30.4 points. GPT-5.5 trended -1.6, slipping from 64.66 to 63.08 points, with a mean of 76.8 points and volatility of 33.5 points.

The declines of Gemini 3.1 Pro and GPT-o3 were accompanied by widening score fluctuations on the execution dimension. After a strong first-day score, Gemini 3.1 Pro incurred consecutive material constraint deductions, pulling its final-day score down by 11.2 points. GPT-o3 posted a mean of 87.1 points but closed at just 69.56 points, indicating unstable engineering judgment side scores across continuous testing.

Flat Models and High-Volatility Signals

GLM-4.6 trended +1, from 87.54 to 88.58 points, with a mean of 71.3 points and volatility of 43.6 points. Doubao Pro trended -0.7, from 86.54 to 85.89 points, with a mean of 87.1 points and volatility of 9.7 points.

GLM-4.6's volatility of 43.6 points is the highest among all models, with large standard deviations in both the execution and material constraint dimension scores, resulting in low stability. Doubao Pro's volatility of just 9.7 points represents the strongest output consistency.

Integrity Rating Change Analysis

GLM-4.6 and Doubao Pro received a "warn" integrity rating on the first day, which improved to "pass" for the remaining six days. Grok 4 recorded a single "warn" on the fifth day, with "pass" for the rest. Models whose integrity ratings improved from "warn" to "pass" showed a marked reduction in material constraint deductions this week, indicating fewer factual deviations in their outputs.

Practical Implications for Users

Teams with heavy code-execution needs can prioritize Grok 4, whose mean of 90 points and volatility of 13.6 points make it well suited for scenarios requiring stable output. For use cases sensitive to material fidelity, Doubao Pro, with volatility of 9.7 points and an integrity rating that turned to "pass," is a strong fit for long-form text generation tasks. Claude Opus 4.7, closing at 93.08 points, suits prototype development requiring high engineering judgment side scores.

Developers relying on Gemini 3.1 Pro should note its -11.2 trend; execution dimension fluctuations in continuous testing may lead to degraded consistency after deployment. GPT-o3 has a mean of 87.1 points but closed at 69.56 points, so production environments with strict stability requirements need additional validation.

Strategic Assessment

Based on this week's data, Claude Opus 4.7 and Claude Sonnet 4.6 show pronounced upward trends, with mean and final-day scores improving in tandem, warranting continued observation of their primary leaderboard performance next cycle. Gemini 3.1 Pro and GPT-o3 are trending downward with volatility exceeding 30 points, signaling reduced consistency in Smoke quick tests. GLM-4.6's high volatility of 43.6 points coexists with an improved integrity rating, requiring verification of whether its execution dimension has entered a stable range. Grok 4, with a mean of 90 points and the lowest volatility, is supported by current data as being undervalued across the two primary leaderboard dimensions.


Data source: YZ Index | Run #300 | View Raw Data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!