Claude Opus 4.7 Scores 72 in Code Execution, 85 in Material Constraints; Smoke Evaluation Main Leaderboard Slips 2.4 Points

In today's Smoke evaluation, Claude Opus 4.7's code execution dimension score fell from yesterday's 100.00 to 72.00, while the material constraints dimension rose from 56.00 to 85.00, and the overall main leaderboard dropped from 80.20 to 77.85.

Data Facts: Opposite Movements in Two Core Dimensions

The Smoke evaluation contains only 10 questions per day, with 2 questions per dimension. Today, the code execution dimension scored 72.00, down 28 points from yesterday; the material constraints dimension scored 85.00, up 29 points from yesterday. Engineering judgment rose from 94.50 to 100.00, while task expression fell from 66.70 to 58.90. The integrity rating remained pass.

The main leaderboard is weighted from code execution and material constraints. Today's 77.85 is down 2.4 points from yesterday. The single-day drop in code execution was much larger than the main leaderboard drop, while the single-day rise in material constraints fully offset part of the loss.

Causes: Question Sampling Volatility Is More Likely

The Smoke evaluation has only 2 questions per dimension, so each question carries extremely high score weight. The code execution dimension scored a perfect 100.00 yesterday and 72.00 today; the 28-point difference most likely comes from differences in question difficulty sampling. The material constraints dimension was low at 56.00 yesterday and 85.00 today, also pointing to a score rebound caused by question sampling.

If the model had truly degraded, it would usually affect multiple dimensions at the same time. Today, engineering judgment instead rose to a perfect 100.00, and although task expression fell by 7.8 points, the decline was far smaller than that of code execution. The two main leaderboard dimensions moved in completely opposite directions, further supporting question sampling rather than a systemic decline in model capability.

Implications for Users

Teams that rely heavily on code execution should allow a buffer for single-day score fluctuations when choosing Claude Opus 4.7. In the Smoke evaluation, a code execution score of 72.00 is still in the usable range, but if similar low scores appear for multiple consecutive days, the model's stability in complex code tasks should be reassessed.

For scenarios sensitive to material constraints, today's 85.00 shows that Claude Opus 4.7 has improved in following given materials. Developers can prioritize testing the model's performance today in tasks requiring strict material fidelity.

Strategic Judgment: No Need to Lower Priority for Now, but Next-Day Verification Is Needed

Based on today's data, Claude Opus 4.7's main leaderboard dropped only 2.4 points, and there is no signal of systemic degradation in core capabilities. The sharp opposite fluctuations in code execution and material constraints are more consistent with the statistical characteristics of 2-question sampling.

If the code execution dimension rebounds to above 90 in tomorrow's Smoke evaluation, today's result can be confirmed as normal fluctuation; if it remains below 80, the code execution dimension should be placed on the key watch list. Current data does not support a negative strategic adjustment to Claude Opus 4.7's overall capabilities.


Data source: YZ Index | Run #345 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!