Claude Sonnet 4.6 Code Execution Drops 22 Points, Material Compliance Rises 25.7 Points

Claude Sonnet 4.6 saw its code execution dimension score drop from 97.00 points yesterday to 75.00 points in today’s Smoke evaluation, a decline of 22 points.

Score Comparison Facts

The material compliance dimension score rose from 60.20 to 85.90 points, an increase of 25.7 points. Engineering judgment (side benchmark, AI-assisted evaluation) dropped from 94.50 to 75.00 points, and task expression (side benchmark, AI-assisted evaluation) fell from 95.00 to 90.00 points. The main benchmark score edged down from 80.44 to 79.91 points, a decline of only 0.5 points. The integrity rating remained at pass.

Analysis of Fluctuation Causes

The Smoke evaluation consists of only 10 questions per day, with 2 questions per dimension, so single-day score variations are naturally large. The sharp opposite swings in the code execution and material compliance dimensions—two main benchmark dimensions—are most likely due to sample variance from question selection rather than genuine model capability degradation. If a code execution question involves complex multi-step reasoning or edge cases, a single model mistake can depress the score by 22 points; conversely, if a material compliance question shifts toward simple fact-checking, the model can achieve a high score. The engineering judgment side benchmark also fell by 19.5 points, further indicating that reasoning-heavy questions were disproportionately weighted in this draw.

Implications for Users

Teams heavily reliant on code execution should perform additional manual verification of Claude Sonnet 4.6’s outputs today and in the coming days, especially in multi-file refactoring and algorithm optimization scenarios. The significant increase in material compliance scores suggests improved performance in strictly following user instructions and avoiding hallucinations, which is more favorable for document generation and contract review scenarios that require high-fidelity output. The main benchmark dropped only 0.5 points, indicating that overall capability remains within a usable range, but a 22-point swing in a single dimension exceeds the stability threshold acceptable to most enterprises.

Strategic Assessment

Current data does not support the conclusion of “model degradation,” because the simultaneous surge in material compliance and the plunge in code execution mirror each other—typical characteristics of question variance rather than systematic capability decline. It is recommended that the next Smoke evaluation focus on whether the code execution dimension rebounds above 90 points; if it remains around 75 points for two consecutive days, a longer-term stability verification should be initiated. The 19.5-point drop in the engineering judgment side benchmark is also worth parallel tracking to rule out potential consistency issues in the model’s complex reasoning chains.

For enterprises undergoing model selection, Claude Sonnet 4.6 can still be the first choice for scenarios with high material compliance requirements, but code-execution-intensive projects should prepare backup models or increase test case coverage. For developers relying on this model for automated code generation, today’s data suggests raising unit test ratios to hedge against single-dimension fluctuation risks.


Data source: YZ Index | Run #250 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!