Claude Sonnet 4.6 Code Execution Plunges 27.4 Points, While Material Constraints Soar 37.7 Points

In today's Smoke evaluation, Claude Sonnet 4.6's Code Execution dimension score fell from yesterday's 71.90 points to 44.50 points, a drop of 27.4 points; the Material Constraints dimension rose from 50.60 points to 88.30 points, an increase of 37.7 points; the main leaderboard score rose from 62.32 points to 64.21 points, an increase of only 1.9 points.

Score Comparison Data

Engineering Judgment (side leaderboard, AI-assisted evaluation) rose from 75.00 points to 100.00 points, while Task Expression (side leaderboard, AI-assisted evaluation) fell from 83.90 points to 45.00 points. The integrity rating remains pass. The Smoke evaluation includes only 10 questions per day, 2 per dimension, so single-day score fluctuations are within the normal range, but this time the sharp changes in opposite directions for Code Execution and Material Constraints exceeded previous single-day records.

Cause Analysis

The Code Execution dimension has only 2 questions. Yesterday it may have drawn questions requiring multi-step debugging or complex logic, while today it drew simple calculations or known-pattern questions, causing the score to drop 27.4 points directly. In the Material Constraints dimension, yesterday it may have encountered constraint tasks requiring strict fidelity to the source text, while today it drew looser materials, making model outputs easier to pass scoring. The Task Expression dimension is likewise limited to 2 questions; yesterday's high score may have come from clearly structured outputs, while today's low score may have come from drawing tasks requiring precise formatting. The main leaderboard score is a weighted average of Code Execution and Material Constraints. After one dimension fell and the other rose, the net gain was only +1.9 points, indicating the main leaderboard has some buffer against extreme volatility.

Implications for Users

Teams that emphasize code execution and rely on Claude Sonnet 4.6 for algorithm or data processing tasks should add a manual review step, because a single-day level of 44.50 points may mean more execution errors in real-world scenarios. For scenarios sensitive to material fidelity, such as contract review or technical document generation, today's 88.30-point performance shows the model has potential for constraint adherence and can be prioritized for testing in such tasks. Engineering Judgment rose to a perfect score, meaning that under the side leaderboard evaluation the model's outputs on engineering decision-type questions are more reliable, but Task Expression's 45 points suggest possible formatting or logical jumps in scenarios requiring clear expression.

Strategic Judgment

This change is most likely caused by question-sampling fluctuation rather than genuine model degradation. The Smoke evaluation sample size is too small, with only 2 questions per dimension; any slight difference in the difficulty or constraint strength of a single question can amplify the score gap. The main leaderboard score remains in the 64-point range, indicating no systemic decline in overall capability. For the next evaluation, it is recommended to focus on whether the Code Execution dimension rebounds above 60 points and whether Material Constraints falls back below 70 points. If both dimensions simultaneously remain at their current extreme levels, further verification of the model's stability on the corresponding tasks will be needed.

For enterprises currently selecting a model, Claude Sonnet 4.6's single-day high score in the Material Constraints dimension can serve as a reference, but the low Code Execution score suggests it should not be directly deployed as a core code generation tool. Developers who rely on this model for multi-round debugging tasks should prepare backup models or manual review processes to cope with possible score fluctuations.


Data source: YZ Index | Run #361 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!