Claude Opus 4.7 fell from 96.01 to 82.16 on the main leaderboard in today's Smoke benchmark.
Data Facts: Main Leaderboard and Sub-Dimension Score Changes
The code execution dimension dropped from 97.00 yesterday to 75.00 today, a decline of 22 points; material constraints fell from 94.80 to 90.90, a decline of 3.9 points; engineering judgment held steady at 100.00; task expression rose from 50.00 to 80.00, an increase of 30 points. The main leaderboard declined 13.9 points overall. The integrity rating remained at pass.
Cause Analysis: Sampling Fluctuation or Genuine Degradation
The Smoke benchmark uses only 2 questions per dimension each day, an extremely small sample size, so single-day score fluctuations fall within the normal range. The 22-point drop in the code execution dimension most likely stems from the two questions drawn today being harder than yesterday's average, or from localized errors by the model in specific coding scenarios. Material constraints fell only 3.9 points, a far smaller margin than code execution, indicating no systemic decline in the model's underlying ability to stay faithful to source material. Task expression rose 30 points in the opposite direction, further confirming the randomness of today's question sampling: the same model can show score differences of dozens of points across different question combinations.
Data source: YZ Index | Run #341 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接