In today's Smoke evaluation, GPT-6.1 Sol's main leaderboard score fell from 86.25 to 77.07, with the code execution dimension dropping from 75.00 to 58.30 — a decline of 16.7 points.
Score Comparison Data
Yesterday: code execution 75.00, material constraints 100.00, engineering judgment 100.00, task expression 72.50, main leaderboard 86.25. Today: code execution 58.30, material constraints 100.00, engineering judgment 100.00, task expression 87.50, main leaderboard 77.07. The integrity rating remains pass.
Breaking Down the Data
The 9.2-point drop on the main leaderboard was driven entirely by the code execution dimension; material constraints and engineering judgment showed zero change, while task expression actually rose by 15 points. The Smoke evaluation contains only 10 questions per day, 2 per dimension, so each individual question carries a high score weight — a miss on 2 questions is enough to cause large fluctuations in a dimension's score.
Cause Analysis
Today's 58.30 in the code execution dimension most likely stems from the 2 questions drawn being of a different difficulty or type than yesterday's. Material constraints and engineering judgment holding at full marks indicates no systematic degradation in the model's underlying capabilities. The rise in task expression suggests the model adapted better to today's questions on the expression dimension. The inherent variance of a 10-question single-day test is sufficient to explain the 16.7-point drop, with no need to assume a real decline in model capability.
Implications for Users
Development teams that rely heavily on code execution should add a manual review step on days when the Smoke evaluation fluctuates, rather than directly adopting code snippets generated by the model. A model that maintains full marks on material constraints can still be used for scenarios requiring strict fidelity to source text. The rise in task expression offers a positive signal for product documentation generation scenarios that require clear output.
Strategic Assessment
This main leaderboard decline falls within the normal range of single-day fluctuation for the Smoke evaluation; the 16.7-point drop in the code execution dimension was not accompanied by simultaneous declines in other dimensions, so it does not yet constitute a degradation signal requiring immediate attention. If the code execution dimension remains below 65 in the next evaluation, it will need to be verified whether the model has entered a sustained low range.
For enterprise model selection, the historical median of GPT-6.1 Sol in the code execution dimension remains higher than most competitors, and single-day data should not change the overall assessment. Developers can continue using the model for non-critical code tasks, but for code generation in production environments, a manual verification process is recommended.
The material constraints dimension holding at full marks demonstrates that the model maintains stable output in constraint adherence. The engineering judgment dimension likewise holds at 100.00, making it suitable for scenarios requiring engineering decision support. The task expression dimension rose from 72.50 to 87.50, showing the model has immediate adaptability on expression tasks.
Taking all of today's dimension data together, GPT-6.1 Sol showed no cross-dimensional systematic decline; only the code execution dimension produced relatively large variance, limited by the number of test questions. The next round of Smoke evaluation results will provide a clearer basis for judging the nature of the fluctuation.
Data source: YZ Index | Run #358 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接