The seven-day Smoke evaluation from July 13 to July 19, 2026 shows that Claude Opus 4.7 ranks first with an average score of 86.9, while GPT-o3 dropped from 97.36 on the first day to 66.86 on the last day, a trend decline of 30.5 points.
Rising Models: Claude Series and Qwen3 Max
Claude Opus 4.7 scored 82.95 on the first day and 95.19 on the last day, with a trend increase of 12.2 points and a volatility of 20.9, making it the only rising model with volatility below 25 this week. Claude Sonnet 4.6 rose from 79.6 to 90.02, a trend increase of 10.4 points, with an average of 73.1. Qwen3 Max rose from 72 to 82.23, a trend increase of 10.2 points, with a volatility of 19.8, the smallest volatility among all models.
The above three models remained relatively stable in the two main benchmark dimensions: code execution and material constraints. Claude Opus 4.7 has not received an integrity warn for seven consecutive days and has the smallest score fluctuation in the execution dimension, indicating its highest consistency in answering the 10 daily questions of Smoke.
Declining Models: GPT-o3 and Doubao Pro See the Largest Drops
GPT-o3 scored 97.36 on the first day and 66.86 on the last day, a trend decline of 30.5 points, with an average of 71.1 and a volatility of 97.4, making it the model with the largest volatility among all models. Doubao Pro fell from 96.7 to 49.21, a trend decline of 47.5 points, with an average of 72.9 and a volatility of 50.1. Gemini 3.1 Pro dropped 40.3 points, and Gemini 2.5 Pro dropped 23.5 points.
The decline of these models is mainly concentrated in the code execution dimension. GPT-o3 had significantly low scores on day 4 and day 6, directly pulling down the weekly average. Doubao Pro lost points multiple times in the material constraint dimension, causing its last-day score to fall below 50.
Volatility Analysis: GPT-5.5 and GLM-4.6 Show Worst Consistency
GPT-5.5 has a volatility of 100 points, with a first-day score of 81.96 and last-day score of 81.44, appearing flat but having the largest standard deviation of daily scores. GLM-4.6 has a volatility of 53.2 points, dropping from 66.78 to 31.44, with an average of only 58.2. DeepSeek V4 Pro has a volatility of 32.7 points, a decline of 17.1 points.
High volatility means significant differences in the model's answers to similar questions. GPT-5.5 can score as high as 95 on one day in the execution dimension and drop to 40 the next day. This inconsistency poses a risk for scenarios requiring stable output.
Separate Integrity Rating Observation
Models with a warn this week include GPT-o3, Grok 4, Doubao Pro, GPT-5.5, Claude Sonnet 4.6, and GLM-4.6. Grok 4 and GLM-4.6 each received two warns, and GLM-4.6 still had a warn on the last day. Claude Opus 4.7 and Qwen3 Max passed the entire week with no warn records.
An integrity warn is typically associated with hallucinations in responses or deviations from material constraints. Models with warns are more prone to lose points in the material constraint dimension in the Smoke evaluation.
Specific Implications for Users
Teams that prioritize code execution should give priority to Claude Opus 4.7, whose average score of 86.9 and volatility of 20.9 provide high consistency. Scenarios that rely on material faithfulness should avoid GPT-o3 and Doubao Pro, as both have last-day scores below 70.
For production environments with high stability requirements, Qwen3 Max's volatility of 19.8 and full-pass integrity record are the best choice currently. GLM-4.6, with an average of 58.2 and multiple warns, is not suitable for tasks requiring continuous stable output.
Strategic Assessment
Claude Opus 4.7 has the best data this week, with the highest average and smallest volatility, making it suitable as a benchmark model. GPT-o3 and Doubao Pro have dropped more than 30 points and need to be verified in the next Smoke data to determine whether this is a temporary fluctuation or a sustained decline.
Although GPT-5.5 has an average of 72, its volatility of 100 makes its actual usability lower than what the average suggests. Qwen3 Max rose by 10.2 points with the smallest volatility, and its execution dimension score changes are worth continuous tracking.
Data source: YZ Index | Run #237 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接