Claude Opus 4.7 scored 83.24 points on today's Smoke evaluation main leaderboard, down 10.3 points from yesterday's 93.54. The core reason is that the code execution dimension fell from 97.00 to 75.00 points.
Score Breakdown and Direct Evidence
The main leaderboard is composed solely of two weighted dimensions: code execution and material adherence. Yesterday, code execution scored 97.00 and material adherence scored 89.30; today, code execution scored 75.00 and material adherence scored 93.30. The 22-point drop in the code execution dimension alone directly pulled the main leaderboard down by 10.3 points, while material adherence actually rose by 4 points, partially offsetting the decline. Engineering judgment remained unchanged at 100.00 points, task expression rose from 78.90 to 90.00 points, and the integrity rating remained "pass."
Cause Analysis: Draw Variance or Capability Degradation
The Smoke evaluation runs only 2 questions per dimension daily, for a total of 10 questions. Losing points on just 2 questions in the code execution dimension can produce a single-day fluctuation of 22 points, which is directly tied to the test scale. The material adherence dimension rose during the same period, indicating that the model's overall state did not decline across the same batch of questions. Existing data only shows abnormal performance on 2 code execution questions and cannot support a conclusion of "genuine model degradation"; it is more likely random fluctuation from question draw.
Specific Implications for Users
Teams that heavily rely on code execution should add local validation steps after today's test, as 75.00 points means at least one of the two questions in this dimension showed a clear error. Developers who saw material adherence scores rise can continue to trust Claude Opus 4.7's material fidelity. The task expression dimension rose to 90.00 points, which has minimal impact on scenarios requiring structured output.
Strategic Assessment
A single-day code execution fluctuation of 22 points is within the normal range for a 10-question test and does not constitute a signal warranting continued attention. If the code execution dimension remains around 75 points in the next evaluation round, multi-day continuous tracking should be initiated. For now, based on a single day of data, Claude Opus 4.7's main leaderboard performance remains within the normal fluctuation range.
Data source: YZ Index | Run #299 | View Raw Data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接