In today's Smoke evaluation, Qwen3 Max's code execution score dropped from 96.70 to 75.00, its material constraint score rose from 48.80 to 69.20, and its main leaderboard score fell from 75.15 to 72.39.
Data Facts: Dimension Score Comparison
The code execution dimension changed by -21.7 points, the material constraint dimension changed by +20.4 points, engineering judgment rose from 36.10 to 47.20, task expression fell from 60.00 to 41.70, and the integrity rating shifted from pass to warn. The main leaderboard fell by only 2.8 points.
Cause Analysis: Draw Fluctuation or Real Degradation
The Smoke evaluation includes only 10 questions per day, 2 per dimension, making the sample size extremely small. The simultaneous sharp declines in code execution and task expression, alongside the rises in material constraint and engineering judgment, indicate that difficulty distribution shifts caused by question draws are the primary driver. The main leaderboard score dipped only 2.8 points, showing that the opposing movements of the two core dimensions offset each other and have not yet formed a consistent degradation signal.
The shift of the integrity rating to warn is the only point in this change requiring extra attention. This rating directly affects model admission eligibility, and its trigger mechanism is directly linked to the low scores in code execution and task expression. However, it currently sits in the warn range, not fail.
Implications for Users
Teams that heavily depend on code execution scenarios should add manual review steps after today. With code execution at 75.00, now below material constraint at 69.20, the model's stability has declined for tasks requiring precise computation or structured output. The rise in material constraint to 69.20 positively affects scenarios requiring strict adherence to input constraints.
The gap between engineering judgment at 47.20 and task expression at 41.70 shows that the model performs relatively better on engineering decision problems than on open-ended expression tasks. Developers relying on task expression should temporarily lower their expectations for Qwen3 Max.
Strategic Assessment
This change is most likely caused by question draw fluctuation rather than model capability degradation. The main leaderboard fell only 2.8 points with the two core dimensions moving in opposite directions, consistent with the normal range for small-sample rapid evaluations. The warn integrity rating is the only signal requiring verification in the next round; if code execution and task expression remain at low levels in the next evaluation, real consistency issues with the model should be considered.
Current data does not support judging Qwen3 Max as overvalued or undervalued; it merely indicates that the single-day fluctuation has touched the integrity threshold. It is recommended that the next Smoke evaluation focus on the recovery of the code execution and integrity rating metrics.
Data source: YZ Index | Run #309 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接