In today's Smoke evaluation, Claude Sonnet 4.6's material constraint score fell from 100.00 yesterday to 84.20, a drop of 15.8 points, while its code execution score rose from 50.00 to 97.00; its overall main leaderboard score rose from 72.50 to 91.24.
Data Breakdown: Sharp Opposite Swings in Two Core Dimensions
The Smoke evaluation has only 2 questions per dimension each day, so the sample size is extremely small. A single-day loss of 15.8 points in material constraint means at least one question had a clear deduction in citation fidelity or constraint adherence; a single-day increase of 47 points in code execution indicates that another question or both questions improved substantially in execution accuracy. The main leaderboard score of 91.24 is weighted from code execution and material constraint, and the huge rebound on the execution side masked the decline on the constraint side.
Cause Analysis: Random Draw Fluctuation or Real Degradation
Looking at the score range, both dimensions simultaneously showing single-day changes of more than 15 points is consistent with the typical characteristics of a small-sample random draw. Material constraint scored a perfect 100.00 yesterday and 84.20 today; the difference may come from just one question that required answering strictly according to the material. Code execution scored 50.00 yesterday and 97.00 today, which may similarly come from one execution-type question being easier or the model happening to hit the correct path. With no consecutive multi-day same-direction trend data to support it, this is more likely attributable to question randomness rather than systematic degradation of model capability.
Specific Implications for Users
For scenarios that rely heavily on material fidelity (such as legal contract extraction, product specification verification, and internal enterprise knowledge base Q&A), Claude Sonnet 4.6's material constraint performance of 84.20 today means the probability of error in a single call has increased. It is recommended to add manual review or multi-model cross-validation in critical processes.
For teams with strong code execution needs (such as automated script generation and data processing pipelines), today's score of 97.00 shows that the model has reached a relatively high usable level on execution-type tasks, and its usage scope can be expanded in controlled environments.
Strategic Judgment
Based on single-day data, the 15.8-point drop in material constraint does not yet constitute a model degradation signal; it is more likely inherent fluctuation in the Smoke evaluation. The integrity rating remains at pass, with no integrity issues appearing. It is recommended to continue tracking the material constraint score next period, and if it is below 90 for two consecutive days, then initiate an in-depth retest; currently there is no need to downgrade the assessment of Claude Sonnet 4.6's overall capability.
Main leaderboard 91.24 vs yesterday's 72.50; the execution-side rebound masks the constraint-side decline, and in the short term observation remains the main approach.
Data source: YZ Index | Run #334 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接