Qwen3 Max Main Score Plunges 14.9 Points, Code Execution Drops from 96.9 to 65.6

Qwen3 Max Main Score Plunges 14.9 Points, Code Execution Drops from 96.9 to 65.6

Qwen3 Max's main score in today's Smoke evaluation fell from 82.23 points yesterday to 67.31 points, a drop of 14.9 points, and the code execution dimension dropped from 96.90 points to 65.60 points.

Score Change Breakdown

The code execution dimension decreased by 31.3 points in a single day, the material constraint dimension rose from 64.30 points to 69.40 points, the engineering judgment (side ranking, AI-assisted evaluation) dropped from 36.30 points to 11.30 points, the task expression (side ranking, AI-assisted evaluation) rose from 46.70 points to 48.90 points, and the integrity rating changed from pass to warn.

Possible Cause Analysis

The Smoke evaluation has only 10 questions per day, with 2 questions per dimension. Daily draw fluctuations are normally within an acceptable range. However, the 31.3-point drop in the code execution dimension far exceeds the typical fluctuation range, and the engineering judgment (side ranking, AI-assisted evaluation) simultaneously saw a 25-point decline, indicating that the model experienced systematic failures in questions requiring multi-step reasoning and code generation. The material constraint dimension actually rose by 5.1 points, suggesting that the model still maintains basic fidelity when referencing given materials, with the issues concentrated on the stability of code generation and logical chains.

The integrity rating changed from pass to warn, further pointing to factual deviations or instruction divergence in some questions, which aligns with the sharp drop in the code execution dimension. The available data cannot determine whether the question draw happened to hit the model's weak points, or whether the model's true capability has degraded, but the 31.3-point single-dimension decline exceeds what a normal draw explanation can account for.

Implications for Users

Development teams heavily reliant on code generation should immediately reduce their trust in Qwen3 Max. The code execution dimension dropping from 96.90 points to 65.60 points means that code outputs that were previously usable now have a high probability of logical errors or runtime failures, requiring additional manual review.

For RAG scenarios that demand high fidelity to source materials, the impact is smaller, as the material constraint dimension remains at 69.40 points, with relatively stable performance in referencing given text. The risk is highest for complex task scenarios that require both code execution and engineering judgment.

Strategic Assessment

The 14.9-point main score drop this time is directly linked to the simultaneous sharp declines in both the code execution and engineering judgment dimensions. The minor changes in material constraint and task expression cannot offset the core capability loss. Qwen3 Max's consistency in Smoke evaluation has shown clear issues. It is recommended to focus on tracking whether the code execution dimension recovers in the next evaluation cycle. If this dimension continues to stay below 80 points, consideration should be given to removing it from the primary model list for code generation tasks.

The conclusion supported by available data is: Qwen3 Max's performance today has triggered a concern threshold, with the abnormal decline in code execution capability and the change in integrity rating together constituting a clear risk signal.


Data source: YZ Index (YZ Index) | Run #238 | View Raw Data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!