In today's Smoke evaluation, Claude Opus 4.7's Material Constraints score dropped from 100.00 yesterday to 84.10, a decrease of 15.9 points, while Code Execution rose from 75.00 to 98.50, an increase of 23.5 points. The Main Leaderboard score rose from 86.25 to 92.02, an increase of 5.8 points.
Core Dimension Score Comparison
Code Execution and Material Constraints make up the Main Leaderboard. Code Execution scored 98.50 today and Material Constraints 84.10; after weighting the two, the Main Leaderboard score is 92.02. Engineering Judgment rose from 80.60 to 100.00, while Task Expression fell from 84.70 to 45.60. The Integrity rating remains pass.
Possible Causes of Score Changes
The Smoke evaluation has only 10 questions per day, 2 per dimension. The single-day drop of 15.9 points in Material Constraints most likely stems from question sampling fluctuation. With a 2-question sample, one high-difficulty material fidelity question can pull down the dimension's average score. Code Execution rose 23.5 points over the same period, which likewise points to random question draw rather than overall model degradation. Task Expression fell by a larger 39.1 points, but this dimension belongs to the Side Leaderboard and does not directly affect the Main Leaderboard ranking.
If this were genuine model degradation, Code Execution and Material Constraints would usually move in the same direction, whereas today they moved in opposite directions, so question sampling fluctuation has stronger explanatory power. The daily 10-question test design itself allows for a relatively large single-day standard deviation; the 31.7-point stability score already indicates that such fluctuation is within the normal range.
Practical Implications for Users
Teams that emphasize Code Execution can continue using Claude Opus 4.7; today's 98.50 shows that it has reached a high level in this dimension. Scenarios that rely on material fidelity, such as contract extraction, literature citation, and data alignment tasks, require additional manual verification; 84.10 is a clear gap from yesterday's 100.00.
Engineering Judgment scored 100.00 today, which benefits developers who need architectural decisions. Task Expression at 45.60, however, suggests that output quality may decline in scenarios requiring long instruction following or high style consistency.
Strategic Assessment
The Main Leaderboard score of 92.02 is higher than yesterday, and the single-dimension plunge in Material Constraints did not change the overall ranking. Based on the existing score comparison, Claude Opus 4.7's improvement in Code Execution offset the loss in Material Constraints, and its current Main Leaderboard performance remains solid. The next Smoke evaluation should focus on verifying whether Material Constraints rebounds above 95. If it stays below 90 for two consecutive days, it may need to be considered for the watch list.
The Side Leaderboard data for Engineering Judgment and Task Expression is for reference only and does not change the Main Leaderboard conclusion. When enterprises select models, they should prioritize the two auditable indicators, Code Execution and Material Constraints; today's data supports continued observation rather than immediate replacement.
Data source: YZ Index | Run #370 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接