In today's Smoke evaluation, Claude Opus 4.7's code execution score fell from 100.00 to 69.50, its material constraint score rose from 71.90 to 95.00, and its main leaderboard score dropped from 87.36 to 80.98.
Score Change Facts
The code execution dimension changed by -30.5 points, the material constraint dimension by +23.1 points, the engineering judgment dimension by +25 points, and the task expression dimension by +5 points. The main leaderboard fell by 6.4 points overall. The integrity rating remains "pass".
Analysis of Fluctuation Causes
The Smoke evaluation includes only 10 questions per day, with 2 questions per dimension. A single mistake in the code execution dimension can cause a drop of more than 30 points. The simultaneous significant rise in material constraint and engineering judgment indicates that the model's overall output capability has not undergone systematic degradation. The randomness of question sampling is the most likely mechanism behind this change.
Code execution: 69.50 vs. yesterday's 100.00; material constraint: 95.00 vs. yesterday's 71.90. The opposing movement of the two dimensions points to sampling rather than model degradation.
Implications for Users
Developer teams that heavily rely on code execution should wait for several consecutive days of stable Smoke data before deciding whether to adjust their workflows. The improved material constraint score is a positive signal for scenarios that require strict adherence to prompts or formats. Engineering judgment rose to 100.00, and the side leaderboard (AI-assisted evaluation) shows the model performs more consistently on engineering decision-making problems.
Strategic Assessment
The sharp drop in code execution this time was most likely caused by single-day question sampling rather than genuine model capability degradation. The fact that the main leaderboard only fell by 6.4 points further supports this assessment. It is recommended that the next Smoke evaluation continue tracking the code execution dimension, with an in-depth re-test triggered only if the score remains below 80 for two consecutive days. The current data does not support downgrading the selection rating for Claude Opus 4.7.
Engineering judgment rose from 75.00 to 100.00, and task expression rose from 90.00 to 95.00. The simultaneous improvement of the two side leaderboard (AI-assisted evaluation) dimensions confirms that the model's stability on non-code-execution tasks remains unaffected.
For teams currently conducting model selection, today's data for Claude Opus 4.7 only shows single-dimension sampling fluctuation, which does not constitute sufficient reason to abandon the model. The material constraint score of 95.00 and engineering judgment score of 100.00 are well suited to scenarios with high demands on format fidelity and engineering decision-making.
The stability dimension measures the standard deviation of the model's scores across multiple responses to similar questions. This single-day change of 30.5 points falls within the normal range of small-sample sampling. The main leaderboard score of 80.98 remains within the usable range.
Taking all dimension data into account, the anomaly in today's Smoke evaluation for Claude Opus 4.7 is mainly concentrated on a single code execution question, and the significant improvements in material constraint and engineering judgment offset most of the negative impact. Users can continue using the model as originally planned while monitoring whether the code execution score recovers to above 90 over the next two days.
Data source: YZ Index | Run #266 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接