Claude Opus 4.7 Code Execution Plunges from 97 to 50, Main Leaderboard Falls 18.7 Points

In today's Smoke evaluation, Claude Opus 4.7's main leaderboard score fell from 82.29 yesterday to 63.59, a decline of 18.7 points.

Breaking Down the Data

Score comparisons show code execution falling from 97.00 to 50.00, a drop of 47 points, while material constraints rose from 64.30 to 80.20, a gain of 15.9 points. Engineering judgment fell from 100.00 to 94.50, and task expression from 91.70 to 85.80. The main leaderboard is a weighted composite of only code execution and material constraints, so code execution's 47-point decline directly dragged the overall score down by 18.7 points.

The Smoke evaluation uses only 2 questions per dimension each day, so a single miss can cause violent swings in a dimension's score. Today's 50.00 in code execution means that at least one of the two questions in that dimension fell below the passing threshold.

Causes

The 47-point drop in code execution most likely stems from the difficulty of today's randomly drawn questions or the model's failure to respond to specific code constraints. Material constraints rose 15.9 points over the same period, showing that the model's performance in faithfully citing materials did not degrade in tandem. Engineering judgment and task expression slipped modestly by 5.5 and 5.9 points, far smaller than code execution, indicating the regression is concentrated in the code generation stage.

If this were a genuine capability regression, material constraints would not have posted a 15.9-point gain in the opposite direction. The current data better supports a question-sampling explanation: with only 2 questions per day, sampling variance is large, and a single failure can halve a dimension's score.

What This Means for Users

Development teams that rely heavily on code execution should note that Claude Opus 4.7's 50.00 in code execution today means a higher probability of single-call failure in scenarios requiring precise code generation or debugging. The 80.20 in material constraints indicates that scenarios relying on the model to faithfully handle provided documents or specifications can still maintain relatively high reliability.

When enterprises evaluate model selection, if code execution accounts for more than 60% of their workload, today's data suggests adding a human review step or calling other models in parallel to hedge against single-day volatility.

Strategic Assessment

Given a sample of only 2 questions, the gap in code execution from 97.00 to 50.00 is a high-probability random fluctuation rather than a permanent degradation of model capability. The integrity rating remains at pass, with no new risk signals.

The next evaluation should focus on verifying whether code execution rebounds above 90. If it stays near 50 for two consecutive days, genuine model degradation must be considered. A single day of data is not enough to judge whether the model is overrated or underrated; continued tracking of the three-day average for code execution is recommended.


Data source: YZ Index | Run #320 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!