In today's Smoke evaluation, Claude Sonnet 4.6's code execution dimension fell from yesterday's 97.00 to 75.00, a decline of 22 points; the material constraint dimension rose from 67.60 to 93.30, an increase of 25.7 points; the main leaderboard score changed from 83.77 to 83.24, dropping only 0.5 points.
Data Facts: Extreme Inverse Volatility Driven by Small Samples
The Smoke evaluation uses only 10 questions per day, with 2 questions per dimension. A single-day loss of 22 points in code execution means the model's average score on today's 2 code execution questions was significantly lower than yesterday's. The material constraint dimension gained 25.7 points in the same period, indicating the model performed considerably better on today's 2 material constraint questions than yesterday. Engineering judgment dropped from 100.00 to 94.50, while task expression rose from 70.00 to 95.00. The integrity rating remained "pass."
Cause Analysis: Question Sampling Volatility, Not True Model Degradation
The nearly symmetric inverse movement between code execution and material constraint most likely stems from extreme differences in how well today's 4 sampled questions (2 per dimension) matched the model's capabilities. Smoke evaluation questions change daily, giving each individual question high weight. With a 2-question sample, a single difficult question can pull down the dimension average by 22 points. The simultaneous sharp rise in material constraint further supports the "question compatibility variation" explanation over "model capability decline." If the model were undergoing systematic degradation, it would be unusual to see an equally large improvement in another core dimension on the same day.
The modest 5.5-point dip in engineering judgment and the 25-point rise in task expression also fit the pattern of small-sample random fluctuation. The main leaderboard score dipped only 0.5 points, showing that the dramatic swings in code execution and material constraint largely offset each other — a direct consequence of the 2-questions-per-dimension-per-day design.
Implications for Users
Development teams that depend heavily on code execution should note that a single-day Smoke score of 75.00 does not mean the model's true coding capability has fallen to that level. It is advisable to use larger sample tests during model selection or routine validation to avoid being misled by a single day's 2-question results. For scenarios sensitive to material constraint (such as applications with high long-document fidelity requirements), today's 93.30 score shows the model still has stable performance headroom on the corresponding questions.
Enterprises using Smoke scores as a daily monitoring metric should observe overall leaderboard changes rather than looking at any single dimension in isolation. The inverse swings between code execution and material constraint serve as a reminder of the inherent limitations of small-sample quick testing.
Strategic Assessment
Based on the current single-day data, Claude Sonnet 4.6's code execution dimension should not be immediately judged as true degradation. The main leaderboard score holding at 83.24 indicates no systematic decline in overall capability. The next evaluation cycle needs to verify whether code execution remains below 85 points; only if similarly low scores appear for two consecutive days should the possibility of model capability change be considered. Current signals point more toward normal fluctuation from question sampling.
For developers relying on this model, existing usage strategies can be maintained in the short term, while expanding internal test coverage of code execution tasks to hedge against the randomness of small-sample evaluation.
Data source: YZ Index | Run #299 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接