Claude Opus 4.7 Scores 75.00 in Code Execution, 100.00 in Material Constraint, Main Leaderboard Rises 3.7

In today's Smoke evaluation, Claude Opus 4.7's code execution score fell from 99.60 yesterday to 75.00, while its material constraint score rose from 61.70 to 100.00. The main leaderboard score increased from 82.55 to 86.25.

Direct Facts Behind the Score Changes

The code execution dimension saw a decline of 24.6 points, the engineering judgment dimension dropped 11.1 points, and the task expression dimension held steady at 91.70. The material constraint dimension rose 38.3 points. As a result, the main leaderboard score gained a net 3.7 points. The integrity rating remained at pass.

Distinguishing Sampling Volatility from Model Degradation

The Smoke evaluation includes only 2 code execution problems and 2 material constraint problems per day. With a daily total of 10 problems, differences in problem draws alone are enough to cause fluctuations of more than 20 points. The code execution drop from 99.60 to 75.00 appeared alongside the material constraint jump from 61.70 to 100.00, indicating that the day's problems presented opposite difficulty distributions across the two dimensions, rather than a systematic degradation of the model in any single capability.

The engineering judgment score fell from 86.10 to 75.00, a smaller decline than that of code execution, while task expression showed zero change. This further supports the conclusion that the fluctuation stems from the problems themselves, not from permanent changes to the model's parameters or reasoning chain.

Impact on Code-Heavy Teams

Developers who depend on code execution accuracy will see a 75.00 score today, down from 99.60 yesterday. Fallback plans or manual checks should be prepared for production environments. The perfect 100.00 material constraint score indicates that the model performs steadily in strictly following user-provided materials, making it suitable for document generation or knowledge extraction scenarios that require high-fidelity output.

Strategic Implications for Enterprises in Model Selection

The main leaderboard score of 86.25, higher than yesterday's 82.55, shows that overall capability has not been impaired by the decline in a single dimension. Enterprises that rely on both code execution and material constraint dimensions should expand test case coverage for code execution scenarios rather than broadly downgrading Claude Opus 4.7's priority. The engineering judgment score of 75.00, down from 86.10 yesterday, warrants additional validation for teams that depend on engineering decision support.

Assessing the Stability Signal

The offsetting movements of -24.6 points in code execution and +38.3 points in material constraint fall within the normal range for a daily sample of only 4 problems. Current data does not support a conclusion of genuine model degradation. Confirmation of degradation should wait for the next evaluation round, and would require code execution to persistently stay below 85 or material constraint to fall back below 70.

The simultaneous appearance of a 75.00 code execution score and a 100.00 material constraint score indicates that the day's problem difficulty was complementary across the two dimensions, rather than reflecting a structural decline in model capability.

Based on the current score comparison, Claude Opus 4.7's main leaderboard ranking has improved thanks to the perfect material constraint score, while the single-day low in code execution is more likely a result of sampling volatility. Teams with heavy code execution needs should add verification steps, while the model can continue to be used in scenarios sensitive to material constraint fidelity.


Data source: YZ Index | Run #280 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!