Claude Sonnet 4.6 Smoke Evaluation Main Leaderboard Plunges 17 Points, Code Execution Falls from 100 to 50

In today's Smoke evaluation, Claude Sonnet 4.6's main leaderboard score fell from 89.52 to 72.50, a drop of 17 points, while the code execution dimension plunged directly from 100.00 to 50.00.

Dimension Breakdown: The Direct Source of the Main Leaderboard Decline

The main leaderboard is weighted from only two dimensions: code execution and material constraint. Yesterday, code execution scored 100.00 and material constraint 76.70; today, code execution scored 50.00 and material constraint 100.00. A 50-point drop in the single code execution dimension directly dragged the main leaderboard down by 17 points, and the 23.3-point rebound in material constraint was not enough to fully offset that loss.

Side leaderboard data shows engineering judgment rising from 50.00 to 83.30, while task expression fell from 95.80 to 65.00. The integrity rating remained at pass, with no threshold issues triggered.

Cause Analysis: Two-Question Sampling Variance Is the Most Likely Explanation

The Smoke evaluation uses only 2 questions per dimension each day, an extremely small sample. The most direct mechanism behind code execution falling from full marks to 50 is that one question was scored a complete loss and the other a partial score. Material constraint simultaneously rising to full marks suggests the material-type questions drawn that day were lower in difficulty or constraint requirements than yesterday's. Engineering judgment rising 33.3 points and task expression falling 30.8 points further confirms that difficulty distribution varies unevenly across dimensions, rather than pointing to an overall decline in model capability.

If this were genuine degradation, it would typically be accompanied by simultaneous declines across multiple dimensions, and since material constraint and code execution are both core capabilities, an extreme opposing swing of 50 points in one rising and one falling would be difficult to explain. The most likely cause supported by the current data is therefore question sampling variance.

What This Means for Users

Teams that rely heavily on code execution should add dedicated code generation tests in production environments. Claude Sonnet 4.6's 50-point-scale swings in code execution on the Smoke evaluation mean that results from a single 2-question test cannot be directly extrapolated to large codebases or complex algorithmic tasks.

For scenarios sensitive to material fidelity (such as contract extraction or compliance review), today's 100.00 score can serve as a reference, but it remains necessary to observe whether material constraint stays consistently high over several consecutive days.

The large swings in engineering judgment and task expression suggest that single-day scores have limited reference value in scenarios requiring structured output or multi-turn dialogue.

Strategic Judgment: Continued Tracking of the Next Round of Data Is Needed

Based on the available score comparison, Claude Sonnet 4.6's 17-point main leaderboard decline was driven mainly by extreme volatility in the single code execution dimension rather than simultaneous degradation across multiple dimensions. If code execution rebounds above 80 in the next Smoke evaluation, this episode can be judged a sampling anomaly; if it remains around 50, the model's stability on code tasks will require further verification.

The current data does not support a conclusion of "model degradation," but it does support the observation of "high volatility in a single dimension." Organizations making selections should list Claude Sonnet 4.6's code execution capability as a separate verification item, rather than directly adopting today's main leaderboard score of 72.50 as the final reference.


Data source: YZ Index | Run #332 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!