In today's Smoke evaluation, GPT-6 Luna's main leaderboard fell from 83.19 to 74.82, a decline of 8.4 points. The code execution dimension dropped from 75.00 to 58.30, task expression fell from 90.00 to 77.50, material constraint rose from 93.20 to 95.00, engineering judgment held at 100.00, and the integrity rating remained pass.
Breaking Down the Data
The Smoke evaluation uses only 2 questions per dimension each day, and the main leaderboard is weighted solely from code execution and material constraint. Today's code execution score of 58.30 is down 16.7 points from yesterday's 75.00, directly pulling the main leaderboard down by 8.4 points. Task expression fell 12.5 points but does not count toward the main leaderboard. Material constraint instead rose 1.8 points, and engineering judgment showed zero movement.
Cause Analysis: Sampling Variance or Genuine Degradation
Code execution consists of only 2 questions, so a single missed question can cause swings on the order of 16.7 points. Material constraint rose over the same period, indicating that the model's overall output quality has not suffered a systemic decline. Engineering judgment scored full marks for a second consecutive day, showing that the model still delivers stable output on questions requiring engineering decisions. The drop in task expression may overlap with the demands for clarity inherent in the code execution questions themselves, causing the same batch of questions to affect both dimensions at once. The available data does not support a conclusion of genuine model degradation; the most likely cause is small-sample variance from question sampling.
Code execution at 58.30 vs. material constraint at 95.00 — the same model on the same day showing a 17.7-point gap between dimensions points to question difficulty distribution rather than a capability cliff.
What It Means for Users
Teams that lean heavily on code execution should be wary of misjudging GPT-6 Luna's 58.30 score today based on a single day of sampling. Scenarios that depend on material constraint can continue as before, since 95.00 shows this dimension remains at a high level. Developers with strict requirements for task expression should add a manual verification step before key outputs, because 77.50 is already below yesterday's level.
- Code generation scenarios: today's 58.30 suggests you may encounter high-difficulty test cases; it is advisable to run Smoke 2–3 more times or build your own benchmark for verification.
- Document constraint scenarios: 95.00 is safe to use; material fidelity is unaffected.
- Engineering decision scenarios: 100.00 remains stable, suitable for workflows that require consistent judgment.
Strategic Assessment
Based on a single day of data, GPT-6 Luna's main leaderboard decline is driven mainly by missed points on 2 code execution questions, while material constraint and engineering judgment did not deteriorate in tandem, so there is no need to immediately lower the assessment of its long-term capabilities. If the code execution dimension rebounds above 70 in the next Smoke evaluation, today can be confirmed as a sampling anomaly; if it remains below 65, further investigation into the model's stability on code-type questions is warranted.
Full marks in engineering judgment for two consecutive rounds, together with a pass integrity rating, together form an entry-threshold signal, indicating that the model's auxiliary dimensions beyond core capabilities remain reliable. Users can continue to watch the average code execution score over the next 3 days; if the standard deviation exceeds 15 points, a multi-model parallel verification mechanism should be added in production environments.
The slight 1.8-point rise in material constraint shows the model still has headroom in following constraints, which can serve as a buffer for fluctuations in code execution. Overall, the most reasonable explanation for today's data is small-sample sampling variance rather than model capability degradation, and it is worth continuing to track scores in the same dimension next period to confirm the trend.
Data source: YZ Index | Run #368 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接