Claude Opus 4.7 Material Constraint Plummets 15.9 Points; Main Leaderboard Falls 7.2 Points in a Single Day

In today's Smoke evaluation, Claude Opus 4.7's material constraint score fell from yesterday's 80.20 to 64.30, and the main leaderboard overall dropped from 91.09 to 83.94.

Core Dimension Breakdown

The code execution dimension held steady at 100.00 with zero fluctuation, while the material constraint dimension showed a -15.9 point drop. The main leaderboard consists solely of these two auditable dimensions, so the single-day decline in material constraint directly pulled down the overall ranking. The engineering judgment side leaderboard fell from 100.00 to 75.00, while the task expression side leaderboard rose from 63.90 to 80.60.

Analysis of Causes of Fluctuation

The Smoke evaluation has only 10 questions per day, with 2 questions corresponding to the material constraint dimension. A score of 64.30 means the average score on the two questions in this dimension was significantly lower than yesterday. Possible causes include differences in content difficulty due to random question selection, or temporary inconsistency in the model's material fidelity. Both code execution questions remained at full marks, indicating the model's capability in this area was unaffected by similar fluctuations.

The -25 point drop in the engineering judgment side leaderboard coincided with the decline in material constraint, suggesting the model performed worse in scenarios requiring engineering decisions based on external materials. The +16.7 point rise in the task expression side leaderboard shows the model improved in task description clarity; the opposite directions further support that this fluctuation is concentrated in material handling.

Specific Implications for Users

For teams that rely heavily on material constraints, such as RAG applications requiring the model to strictly cite given documents or codebases, today's data indicates that Claude Opus 4.7 showed a clear deviation in a single-day test. Teams with stable code execution can continue using the model for pure programming tasks, but should add manual review of citation accuracy.

In scenarios sensitive to material fidelity, such as legal contract review and product specification alignment, a score of 64.30 means at least one of every two questions showed a clear deviation. Developers should add a material-constraint review step to their workflows rather than directly adopting model outputs.

Strategic Assessment

Based on the current score comparison, the -15.9 point drop in material constraint is more likely due to question sampling fluctuation under the single-day 10-question framework, rather than genuine model degradation. The zero change in code execution confirms that the model's core capabilities did not decline overall. The integrity rating remains pass, with no access threshold issues triggered.

If material constraint remains below 65 in the next Smoke evaluation, this signal should be treated as preliminary evidence of declining model consistency. Current single-day data is insufficient to determine whether the model is overestimated or underestimated; it is recommended to add Claude Opus 4.7's material constraint performance to the key tracking list for the next period.

The main leaderboard fell from 91.09 to 83.94, mainly driven by material constraint. The fact that code execution remains at full marks shows Claude Opus 4.7 remains competitive in auditable capabilities, but fluctuations in material constraint have had a measurable impact on overall usability.


Data: YZ Index | Run #324 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!