Claude Sonnet 4.6 Material Constraint Plummets 15.8 Points, Code Execution Rebounds 47 Points

In today's Smoke evaluation, Claude Sonnet 4.6's material constraint score fell from 100.00 yesterday to 84.20, a drop of 15.8 points, while its code execution score rose from 50.00 to 97.00; its overall main leaderboard score rose from 72.50 to 91.24.

Data Breakdown: Sharp Opposite Swings in Two Core Dimensions

The Smoke evaluation has only 2 questions per dimension each day, so the sample size is extremely small. A single-day loss of 15.8 points in material constraint means at least one question had a clear deduction in citation fidelity or constraint adherence; a single-day increase of 47 points in code execution indicates that another question or both questions improved substantially in execution accuracy. The main leaderboard score of 91.24 is weighted from code execution and material constraint, and the huge rebound on the execution side masked the decline on the constraint side.

Cause Analysis: Random Draw Fluctuation or Real Degradation

Looking at the score range, both dimensions simultaneously showing single-day changes of more than 15 points is consistent with the typical characteristics of a small-sample random draw. Material constraint scored a perfect 100.00 yesterday and 84.20 today; the difference may come from just one question that required answering strictly according to the material. Code execution scored 50.00 yesterday and 97.00 today, which may similarly come from one execution-type question being easier or the model happening to hit the correct path. With no consecutive multi-day same-direction trend data to support it, this is more likely attributable to question randomness rather than systematic degradation of model capability.

Specific Implications for Users

For scenarios that rely heavily on material fidelity (such as legal contract extraction, product specification verification, and internal enterprise knowledge base Q&A), Claude Sonnet 4.6's material constraint performance of 84.20 today means the probability of error in a single call has increased. It is recommended to add manual review or multi-model cross-validation in critical processes.

For teams with strong code execution needs (such as automated script generation and data processing pipelines), today's score of 97.00 shows that the model has reached a relatively high usable level on execution-type tasks, and its usage scope can be expanded in controlled environments.

Strategic Judgment

Based on single-day data, the 15.8-point drop in material constraint does not yet constitute a model degradation signal; it is more likely inherent fluctuation in the Smoke evaluation. The integrity rating remains at pass, with no integrity issues appearing. It is recommended to continue tracking the material constraint score next period, and if it is below 90 for two consecutive days, then initiate an in-depth retest; currently there is no need to downgrade the assessment of Claude Sonnet 4.6's overall capability.

Main leaderboard 91.24 vs yesterday's 72.50; the execution-side rebound masks the constraint-side decline, and in the short term observation remains the main approach.

Data source: YZ Index | Run #334 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!