In today's Smoke evaluation, Claude Sonnet 4.6's code execution score fell from 94.50 to 75.00, a decline of 19.5 points, while its material constraint score rose from 43.30 to 97.80, and the overall leaderboard score climbed from 71.46 to 85.26.
Data Facts: Extreme Contrast Under a Single-Day Two-Question Sample
This evaluation uses a quick-test setup of 10 questions per day and 2 questions per dimension. The code execution dimension lost 19.5 points in a single day, while the material constraint dimension gained 54.5 points, a numerical gap of 74 points. Engineering judgment dipped slightly from 88.90 to 86.70, and task expression fell from 90.00 to 85.80. The integrity rating remained at pass.
Cause Analysis: Question Sampling Fluctuation Rather Than Model Degradation
The Smoke evaluation's extremely small sample of only 2 questions per day gives each question's score an outsized weight. In this session, the code execution dimension may have drawn question types unfavorable to this model, resulting in the low 75.00 score; the material constraint dimension, conversely, drew highly matched constraint tasks, pushing the score up to 97.80. The declines in engineering judgment and task expression were both within 5 points, falling within the normal range of fluctuation. Existing data does not support a conclusion of genuine model capability degradation; on the contrary, the 13.8-point rise in the overall leaderboard score shows that the material constraint dimension's extremely positive contribution offset the code execution loss.
Implications for Users
Development teams that rely heavily on code execution should note that a single-day 75.00 score may appear during consecutive unfavorable sampling, and it is advisable to add manual review steps in critical code generation tasks. For scenarios sensitive to material fidelity, such as contract review or data extraction, Claude Sonnet 4.6's 97.80 score this session demonstrates high usability on specific constrained tasks. Engineering judgment and task expression scores remain above the 85-point range, so the impact on product integration requiring structured output is limited.
Strategic Assessment
This session's data is most likely driven by sampling fluctuation rather than a systemic model change. The fact that the overall leaderboard score rose indicates that the material constraint dimension's weight has a greater influence on overall ranking in the current evaluation. The signal that needs verification in the next session is: if the code execution score remains below 80 points in three consecutive Smoke evaluations, a dedicated test with a larger dataset should be initiated. Current single-day data is insufficient to determine whether Claude Sonnet 4.6 is overrated or underrated; it is recommended to place this model on a watch list rather than adjust its priority.
Based on this session's score comparison, Claude Sonnet 4.6's performance in the Smoke evaluation is consistent with the high-variance characteristics of small samples, and its overall capability shows no degradation signals requiring immediate attention.
Data source: YZ Index | Run #268 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接