In today's Smoke evaluation, Claude Sonnet 4.6's material constraint score fell from 93.30 to 62.70, a drop of 30.6 points, while its code execution score rose from 50.00 to 75.00, an increase of 25 points. Its main leaderboard score changed from 69.49 to 69.47, a near-zero change.
Core Data Breakdown
The Smoke evaluation has only 10 questions per day, with 2 questions corresponding to each dimension. Material constraint and code execution both belong to the auditable dimensions of the main leaderboard. Yesterday's material constraint score of 93.30 reflected high material fidelity, while today's 62.70 directly pulled down that dimension. Code execution jumped by 25 points from 50.00, partially offsetting the material constraint loss and keeping the main leaderboard at 69.47.
Engineering judgment rose from 88.90 to 100.00, and task expression rose from 75.00 to 90.80, with both side-leaderboard dimensions moving upward in tandem. The integrity rating remained pass for two consecutive days and did not trigger any threshold.
Analysis of Fluctuation Causes
The one-day 30.6-point drop in material constraint most likely stems from random question sampling. The Smoke evaluation changes questions daily, and with a 2-question sample, a single-question miss can cause a 20-30 point fluctuation. Code execution rose sharply over the same period, indicating no systematic degradation in code-related tasks. The two main leaderboard dimensions moved in opposite directions, pointing to differences in question content rather than an overall decline in model capability.
If this were genuine degradation, code execution and material constraint would usually move in the same direction; this opposite movement is more consistent with sampling fluctuation. Engineering judgment and task expression improved at the same time, also supporting that the model's underlying capability was unaffected.
Specific Implications for Users
For scenarios that prioritize material fidelity, such as long-document summarization, contract clause extraction, and code comment generation, additional human review of Claude Sonnet 4.6's output today is needed. Teams that prioritize code execution can continue using the model, as 75.00 is already above yesterday's level.
For mixed tasks that rely on both material constraint and code execution, it is advisable to add intermediate checkpoints to prevent single-dimension fluctuations from affecting the final result. An integrity rating of pass indicates that the model did not refuse to answer or show obvious hallucinations, and its baseline usability remains within a safe range.
Strategic Assessment
The main leaderboard score of 69.47 is negligibly different from yesterday's 69.49, and the one-day material constraint plunge did not change the overall ranking position. If material constraint rebounds above 80 in the next Smoke evaluation, this drop can be attributed to sampling noise; if it remains below 70, it will be necessary to watch whether the model has entered a new fluctuation range.
The current data does not support a conclusion of genuine model degradation. It is advisable to maintain the existing usage strategy and add additional checks only in material-constraint-sensitive scenarios.
Data from: YZ Index | Run #364 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接