In today's Smoke evaluation, Claude Opus 4.7's main leaderboard score fell from 82.29 yesterday to 63.59, a decline of 18.7 points.
Breaking Down the Data
Score comparisons show code execution falling from 97.00 to 50.00, a drop of 47 points, while material constraints rose from 64.30 to 80.20, a gain of 15.9 points. Engineering judgment fell from 100.00 to 94.50, and task expression from 91.70 to 85.80. The main leaderboard is a weighted composite of only code execution and material constraints, so code execution's 47-point decline directly dragged the overall score down by 18.7 points.
The Smoke evaluation uses only 2 questions per dimension each day, so a single miss can cause violent swings in a dimension's score. Today's 50.00 in code execution means that at least one of the two questions in that dimension fell below the passing threshold.
Causes
The 47-point drop in code execution most likely stems from the difficulty of today's randomly drawn questions or the model's failure to respond to specific code constraints. Material constraints rose 15.9 points over the same period, showing that the model's performance in faithfully citing materials did not degrade in tandem. Engineering judgment and task expression slipped modestly by 5.5 and 5.9 points, far smaller than code execution, indicating the regression is concentrated in the code generation stage.
If this were a genuine capability regression, material constraints would not have posted a 15.9-point gain in the opposite direction. The current data better supports a question-sampling explanation: with only 2 questions per day, sampling variance is large, and a single failure can halve a dimension's score.
What This Means for Users
Development teams that rely heavily on code execution should note that Claude Opus 4.7's 50.00 in code execution today means a higher probability of single-call failure in scenarios requiring precise code generation or debugging. The 80.20 in material constraints indicates that scenarios relying on the model to faithfully handle provided documents or specifications can still maintain relatively high reliability.
When enterprises evaluate model selection, if code execution accounts for more than 60% of their workload, today's data suggests adding a human review step or calling other models in parallel to hedge against single-day volatility.
Strategic Assessment
Given a sample of only 2 questions, the gap in code execution from 97.00 to 50.00 is a high-probability random fluctuation rather than a permanent degradation of model capability. The integrity rating remains at pass, with no new risk signals.
The next evaluation should focus on verifying whether code execution rebounds above 90. If it stays near 50 for two consecutive days, genuine model degradation must be considered. A single day of data is not enough to judge whether the model is overrated or underrated; continued tracking of the three-day average for code execution is recommended.
Data source: YZ Index | Run #320 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接