In today's Smoke evaluation, Qwen3 Max's material constraints dimension dropped from 90.90 points yesterday to 74.00 points, a decrease of 16.9 points; the code execution dimension rose from 25.00 points to 89.30 points, an increase of 64.3 points, and the main leaderboard total score rose from 54.66 points to 82.42 points.
Score Details and Dimension Composition
The Smoke evaluation has a fixed 10 questions each day, with 2 questions each for code execution and material constraints. Today, the two material constraints questions scored 74.00 points in total, down 16.9 points from yesterday's 90.90 points; the two code execution questions scored 89.30 points in total, up 64.3 points from yesterday's 25.00 points. Engineering judgment remained unchanged at 36.10 points, while task expression rose from 33.90 points to 38.90 points. The integrity rating remained pass.
Analysis of Fluctuation Causes
The simultaneous large opposite movements in material constraints and code execution are most likely due to question sampling randomness. The Smoke evaluation has only 4 main leaderboard questions per day, so each question carries high weight; drawing questions of different difficulty or type can cause sharp score swings. The strong rebound in code execution from a low level suggests that today's code questions may better match the model's current strengths; the simultaneous decline in material constraints may mean the questions drawn required strict fidelity to the original text or format constraints. The simultaneous opposite changes in the two dimensions point to question variance rather than systematic degradation in model capability.
Specific Implications for Users
For scenarios that rely heavily on material constraints, such as contract extraction, regulatory Q&A, and faithful rewriting of long documents, today's 74.00 points means single-day output may deviate in format or content, and developers need to add manual review steps. Teams that emphasize code execution can take advantage of today's 89.30-point performance to achieve a higher pass rate in script generation and data processing tasks. For enterprises selecting a model, the main leaderboard score of 82.42 represents only a snapshot for the day and should not be used directly as a basis for long-term deployment.
Strategic Judgment
Based on the available single-day data, the most reasonable explanation for the 16.9-point decline in material constraints and the 64.3-point rebound in code execution is question sampling fluctuation, not genuine model degradation. The stability dimension score of 31.7 already indicates that this model has large single-day score variance, and this opposite movement further confirms insufficient consistency. It is recommended to continue tracking in the next period whether material constraints return to the 90-point range; if it remains below 80 points for two consecutive days, then judge it as a capability change. At present, there is no need to downgrade conclusions about Qwen3 Max's main leaderboard capability.
Data: YZ Index | Run #328 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接