Qwen3 Max's main leaderboard score in today's Smoke evaluation dropped from 96.04 to 85.82, a decline of 10.2 points.
Score Breakdown: Material Constraint Dimension Dominates the Decline
The code execution dimension fell from 100.00 to 99.00, a decrease of only 1 point. The material constraint dimension dropped from 91.20 to 69.70, a decline of 21.5 points. Engineering judgment fell from 55.60 to 50.00, down 5.6 points. Task expression rose from 41.70 to 73.90, an increase of 32.2 points. The main leaderboard is weighted solely by code execution and material constraint, so the 21.5-point drop in material constraint directly dragged down the overall score by 10.2 points.
Cause Analysis: Question Sampling or Real Degradation?
The Smoke evaluation covers only 10 questions per day, with 2 questions per dimension, so single-day fluctuations fall within the normal range. The material constraint dimension's decline this time far exceeds that of code execution, suggesting that the material constraint questions drawn today may have included higher-difficulty fidelity requirements or conflicting instructions. Code execution remains at 99.00, indicating no systematic decline in the model's basic computational capabilities. The task expression side leaderboard rose sharply by 32.2 points, further showing that the model was unaffected on expression-type questions, pointing to the issue being concentrated specifically in the material constraint dimension.
The material constraint dimension fell 21.5 points this time, far higher than code execution's 1 point, indicating that the fluctuation mainly came from the material constraint question draw.
Implications for Users
Teams that prioritize code execution can continue using Qwen3 Max, as the code execution dimension remains at 99.00 with limited impact. For scenarios sensitive to material fidelity, such as contract extraction, policy comparison, and long-document summarization, Qwen3 Max's score of 69.70 today requires additional manual verification. The engineering judgment side leaderboard fell to 50.00, which has limited impact on developers who rely on engineering decision support, given the dimension's relatively low weight.
Strategic Assessment
This decline is primarily driven by the single material constraint dimension, and code execution remains near perfect, indicating typical sampling fluctuation rather than overall model degradation. The integrity rating remains "pass," with no integrity concerns identified. It is recommended that the next Smoke evaluation focus on whether material constraint rebounds above 90; if it stays below 80 for two consecutive periods, further validation of the model's consistency on constraint-type tasks will be needed.
Current data does not support interpreting this fluctuation as a permanent decline in model capability. Qwen3 Max remains at a high level in the code execution dimension, making it suitable for code-centric workflows; for material constraint scenarios, additional verification steps are recommended.
Data source: YZ Index | Run #270 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接