In today's Smoke evaluation, Claude Sonnet 4.6's material constraint score dropped from 100.00 yesterday to 66.40, while the main ranking overall fell from 86.25 to 81.86.
Score Comparison and Key Facts
Code execution rose from 75.00 to 94.50 points, task expression rose from 79.70 to 90.00 points, and engineering judgment fell from 100.00 to 82.50 points. The integrity rating remained "pass." All of the above data comes from the same Smoke evaluation; the comparison between yesterday and today is limited to these two single-day results.
Possible Causes of the Material Constraint Plunge
The Smoke evaluation uses only 10 questions per day, with 2 questions per dimension. A single-day -33.6-point swing in material constraint most likely stems from question sampling variance. Two new questions imposed stricter requirements on material fidelity, directly pulling down the dimension's score. Code execution simultaneously rose +19.5 points, indicating that the model's execution capability did not regress on another set of questions. The -17.5-point movement in engineering judgment likewise points to changes in question difficulty rather than a decline in the model's overall capability.
If this were genuine model regression, multiple dimensions would typically decline in unison. However, both code execution and task expression rose notably today, so the probability of true regression is low.
Practical Implications for Users
Scenarios that rely heavily on material constraint—such as legal contract review, internal document rewriting, and tasks with high long-text fidelity requirements—should be held back from large-scale deployment until the next Smoke evaluation results are released. Code execution has reached 94.50 points, so teams with heavy code-execution needs can continue using the model for programming-related work.
Engineering judgment (side ranking, AI-assisted evaluation) has fallen to 82.50 points, which may affect processes requiring complex decision chains. Developers should add manual review at key decision points.
Strategic Assessment
A single-day -33.6-point drop in material constraint is insufficient to conclude that the model's capability has regressed, but the magnitude of fluctuation in this dimension has exceeded the main ranking's average. If material constraint remains below 80 points in the next Smoke evaluation, the model should be downgraded for use in material-sensitive scenarios. Current data only supports the conclusion that "question sampling caused the drop," not a judgment of "persistent regression."
It is recommended to track Claude Sonnet 4.6's material constraint score as a daily monitoring item, and to revisit model selection only if it remains below 75 points for two consecutive days.
Data source: YZ Index | Run #315 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接