Qwen3 Max Code Execution Plunges 24 Points, Main Ranking Falls 7.4 Points

In today's Smoke evaluation, Qwen3 Max's code execution score fell from 94.00 to 70.00, while its main ranking dropped from 84.51 to 77.16.

Score Change Details

The code execution dimension saw a sharp decline of -24 points, while the material constraints dimension rose from 72.90 to 85.90, a gain of 13 points. Engineering judgment climbed from 11.30 to 36.80, and task expression rose from 63.90 to 75.60. The integrity rating remained at pass.

Main Ranking Composition and Sources of Fluctuation

The main ranking is formed by weighting only two dimensions: code execution and material constraints. The -24 points in code execution directly pulled the main ranking down by 7.4 points, while the +13 points in material constraints partially offset this effect. Engineering judgment and task expression belong to the secondary ranking, are AI-assisted assessments, and are not counted toward the main ranking.

Analysis of Possible Causes

The Smoke evaluation includes only 10 questions per day, 2 per dimension, making the sample size small; question-draw fluctuation is the primary possible factor. Changes in the difficulty or type of code execution questions may cause sharp score swings rather than a systematic degradation in the model's actual capabilities. The rise in the material constraints score indicates that the model has not weakened in tandem on constraint adherence.

If this were genuine degradation, code execution and material constraints would typically move in the same direction, but this time the two dimensions moved in opposite directions, supporting the fluctuation explanation. The increases in engineering judgment and task expression likewise show no signal of an overall capability decline.

Implications for Users

Teams that rely heavily on code execution should note that Qwen3 Max's code execution score in the Smoke evaluation has fallen to 70.00, and single-day performance may be below the previous 94.00 level. For scenarios that depend on material fidelity, its new high of 85.90 can be used as a reference.

When selecting a model, developers should take both the code execution score of 70.00 and the material constraints score of 85.90 into consideration, and avoid making judgments based solely on the main ranking score of 77.16.

Strategic Judgment

This -24 point drop in code execution is more likely due to question-draw fluctuation than to genuine model degradation. The main ranking score of 77.16 is still within the usable range, but consistency needs to be verified with the next round of data. The rises in engineering judgment to 36.80 and task expression to 75.60 do not change the core conclusion of the main ranking.

If code execution scores remain below 80 across multiple subsequent rounds, then its applicability in code-intensive scenarios should be reassessed. The current data only supports paying attention to fluctuation, not concluding that capability has degraded.


Data source: YZ Index | Run #354 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!