In today's Smoke evaluation, Claude Sonnet 4.6's material constraint score fell from 100.00 yesterday to 79.50, a single-day drop of 20.5 points, while code execution rose from 73.70 to 100.00, lifting the main leaderboard from 85.54 to 90.78.
Score Facts and Main Leaderboard Composition
The main leaderboard is weighted solely by two dimensions: code execution and material constraint. Today, code execution scored a perfect 100.00, material constraint scored 79.50, and averaging the two yields a main leaderboard of 90.78. Yesterday, when material constraint was at full marks, the main leaderboard stood at 85.54. Today's 20.5-point loss in material constraint was offset by the 26.3-point gain in code execution, still delivering +5.2 on the main leaderboard.
Engineering judgment fell from 100.00 to 60.50, and task expression from 95.80 to 95.00; both belong to the side leaderboard (AI-assisted evaluation) and do not count toward the main leaderboard ranking. The integrity rating remains "pass."
Potential Cause Analysis
The Smoke evaluation uses only 10 questions per day, with 2 per dimension, making question sampling highly random. The single-day 20.5-point loss in material constraint most likely stems from the two material-constraint questions drawn today imposing stricter fidelity requirements on model output, resulting in deductions in citation or constraint adherence. Both code-execution questions received perfect scores, indicating stable performance on the sampled questions in that dimension.
If this were genuine model degradation, multiple dimensions would typically fluctuate in tandem. Today, however, task expression dropped only 0.8 points, and while the side-leaderboard engineering judgment fell 39.5 points, it does not affect the main leaderboard. A sharp swing in a single dimension paired with a perfect reverse score in another core dimension aligns more with sampling variance than with systematic capability decline.
Practical Implications for Users
For scenarios heavily dependent on material constraint—such as long-document summarization, contract clause extraction, and knowledge-base Q&A—Claude Sonnet 4.6's 79.50 score today means outputs may exhibit factual deviations or omissions. Adding a manual verification step is recommended.
For developer teams whose primary needs center on code execution, today's 100.00 score demonstrates high reliability in tasks such as algorithm implementation and debugging scripts, supporting continued use in production environments.
Mixed teams using both types of tasks should note that the model may show uneven performance across dimensions within the same batch of responses. It is advisable to set separate prompts or post-processing rules by task type.
Strategic Assessment
Based on today's data, Claude Sonnet 4.6's 20.5-point decline in material constraint is most likely attributable to question-sampling variance rather than genuine model degradation. The main leaderboard still achieved net growth thanks to the perfect code-execution score, indicating no systemic issues in core capabilities.
Single-day fluctuation in the Smoke evaluation falls within the normal range. We recommend continuing to track whether the material-constraint score rebounds in the next cycle. Only if material constraint remains below 80 for two consecutive days should deeper evaluation be triggered. Current data does not support a downward revision of the model's overall capability.
The sharp drop on the engineering-judgment side leaderboard (AI-assisted evaluation) can serve as a secondary observation indicator but does not change the main-leaderboard conclusion. The integrity rating remains "pass," and response consistency has not triggered the integrity threshold.
Data source: YZ Index | Run #297 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接