GPT-6 Luna Material Constraint Falls 20 Points, Code Execution Rises 41.7 Points; Main Leaderboard Gains 13.9 Points

In today's Smoke Evaluation, GPT-6 Luna's Material Constraint score fell from 95.00 points to 75.00 points, its Code Execution score rose from 58.30 points to 100.00 points, and its Main Leaderboard score rose from 74.82 points to 88.75 points.

Breakdown of the Data

Comparing yesterday's and today's scores, the Code Execution dimension rose by 41.7 points, the Material Constraint dimension fell by 20 points, the Engineering Judgment dimension remained unchanged at 100.00 points, and the Task Expression dimension rose by 10 points to 87.50 points. The Main Leaderboard is weighted from the two auditable dimensions of Code Execution and Material Constraint, for a final net increase of 13.9 points. The Integrity Rating remained at pass and did not trigger the threshold.

The Smoke Evaluation has only 2 questions per dimension per day, 10 questions in total. A single-day 20-point drop in Material Constraint means that at least one question incurred a clear deduction for material fidelity; a single-day 41.7-point rise in Code Execution means that the corresponding question received a full or near-full score for execution correctness.

Cause Analysis: Draw Variance or Model Degradation

The simultaneous large opposite movements in Material Constraint and Code Execution are most likely due to question sampling variance. With a sample size of only 2 questions per day, one Material Constraint question involving complex citations can pull the score down by 20 points, while one question requiring precise code output can push it up by 41.7 points. Engineering Judgment and Task Expression did not decline in sync, further supporting variance rather than systematic degradation.

If this were genuine model degradation, it would usually appear as consistent declines across multiple dimensions; however, in today's data Engineering Judgment remained at a full score, Task Expression rose instead, and the Integrity Rating also remained at pass. Therefore, the 20-point drop in Material Constraint is closer to statistical noise from single-day question sampling.

Specific Implications for Users

Teams focused on code execution can have higher confidence today, as GPT-6 Luna reached 100.00 points on the corresponding questions; scenarios sensitive to material fidelity require additional manual review, because 75.00 points is already below yesterday's 95.00. Developers relying on this model should refer to multi-day average scores rather than single-day results when making selections.

The Engineering Judgment dimension remained stable at 100.00 points, indicating that side leaderboard capabilities were unaffected; Task Expression rose slightly, indicating an improvement in structured output capability. Overall, single-day data has limited impact on production environment selection.

Strategic Judgment

Based on the available score comparison, GPT-6 Luna showed no signal of systematic degradation in this Smoke Evaluation, and the decline in Material Constraint is more likely due to sampling variance. The Main Leaderboard score actually rose by 13.9 points, indicating that the gain in Code Execution covered the loss in Material Constraint.

The stability dimension previously showed 31.7 points, confirming that this model's scores fluctuate considerably on similar questions. It is recommended to continue tracking the day-to-day standard deviation of Material Constraint and Code Execution in the next period; if Material Constraint remains low for two consecutive days, then evaluate whether to adjust weights or add a manual review step.

The current data does not support the conclusion of "model degradation"; the inherent variance of a single-day 10-question quick test can fully explain all changes.


Data source: YZ Index | Run #370 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!