Claude Sonnet 4.6 Code Execution Drops from 100 to 75 Points, Main Leaderboard Falls 5.5 Points

In today's Smoke evaluation, Claude Sonnet 4.6's code execution score dropped from 100.00 to 75.00, a decline of 25 points, and the main leaderboard overall fell from 80.52 back to 75.00.

Score Change Breakdown

The main leaderboard includes only two dimensions: code execution and material constraint. Today's code execution scored 75.00 and material constraint scored 75.00; averaging the two yields a main leaderboard score of 75.00. Yesterday's material constraint was only 56.70, so today's rise of 18.3 points partially offset the decline in code execution. Engineering judgment dropped from 100.00 to 60.50, while task expression rose from 65.80 to 70.00; these two side-leaderboard dimensions are not counted toward the main leaderboard ranking.

Possible Causes

The Smoke evaluation uses only 10 questions per day, with 2 questions mapping to each dimension. The code execution dimension lost more points today, most likely due to difficulty distribution changes from question sampling rather than a systematic regression in model capability. The simultaneously rising material constraint score also points to random allocation of the day's questions across different dimensions, rather than a consistency issue in any single direction. The integrity rating remains at "pass," with no new signs of violations.

What It Means for Users

Development teams that rely heavily on code execution should add a manual review step when using Claude Sonnet 4.6 for complex algorithm or multi-step reasoning tasks. The material constraint score recovering to 75.00 indicates that the model performs consistently in following user-provided contextual materials, making it suitable for document generation or knowledge extraction scenarios that demand high fidelity.

Strategic Assessment

A single-day 25-point fluctuation is within the normal range for a 10-question quick test, and the current data is insufficient to conclude that the model has truly degraded. It is advisable to continue tracking the standard deviation of scores for the same dimension over the next three days; only if code execution consistently falls below 85 should one consider adjusting model selection priorities. The sharp drop in engineering judgment (side leaderboard, AI-assisted evaluation) can serve as a secondary reference signal but does not change the main leaderboard assessment.

Overall, the alignment between Claude Sonnet 4.6's main leaderboard score of 75.00 and its material constraint score of 75.00 today shows that the model has become broadly balanced across the two core dimensions of code and materials. Enterprises making model selections may treat it as a mid-to-high-tier option, but should confirm stability with multi-day data.


Data source: YZ Index | Run #280 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!