GPT-6 Sol Material Constraint Plummets 33.3 Points as Code Execution Rebounds 25 Points and Main Leaderboard Edges Down

In today's Smoke evaluation, GPT-6 Sol's material constraint score fell from 88.30 to 55.00, a drop of 33.3 points, while its main leaderboard score slipped only slightly from 80.99 to 79.75.

Score Comparison Data

Code execution rose from 75.00 to 100.00, an increase of 25 points; material constraint fell from 88.30 to 55.00; engineering judgment remained unchanged at 94.80; task expression dropped from 100.00 to 77.50, a decline of 22.5 points; integrity rating remained pass.

Breaking Down the Data

The main leaderboard is weighted from two dimensions: code execution and material constraint. A perfect code execution score of 100.00 directly lifted the main leaderboard, while the 55.00 material constraint score offset that gain, leaving the overall score down only 1.2 points. Task expression, as a side leaderboard dimension, also fell by 22.5 points, indicating inconsistent responsiveness to instruction fidelity in this 10-question rapid test.

Cause Analysis

The Smoke evaluation includes only 2 questions per dimension each day, so random question selection may cause sharp single-day score swings. The gap between the 55.00 material constraint score and yesterday's 88.30 may be because the questions drawn this time placed higher demands on material citation, and the model failed to maintain consistent output. Code execution rising to 100.00 shows the model performed stably on this set of coding questions, without similar degradation. The opposite movements across the two dimensions are more consistent with question-draw volatility than with an overall decline in model capability.

Implications for Users

In scenarios that rely heavily on material constraints, such as contract review and literature citation generation, enterprises should add an extra manual verification step to avoid output deviation at the 55.00 score level. Developers who depend on code execution can continue using GPT-6 Sol for programming tasks; the 100.00 score in this run provides short-term confidence. The 77.50 task expression score suggests that in multi-turn instruction scenarios, the model may omit instructions, so splitting complex tasks is recommended to reduce risk.

Strategic Judgment

A single-day drop of 33.3 points in material constraint exceeds the normal fluctuation range and warrants close tracking in the next Smoke evaluation. If material constraint remains below 60 for two consecutive days, systematic issues in the model's material fidelity dimension should be considered. The current main leaderboard score of 79.75 is still in a usable range, but teams sensitive to material fidelity should prepare backup models. Engineering judgment remains stable at 94.80 and can serve as a reference anchor for the model's side leaderboard capabilities.

This evaluation shows a clear divergence between GPT-6 Sol's two core dimensions: code execution and material constraint. The perfect code execution score may stem from a good fit with this run's question difficulty, while the 55.00 material constraint score exposes insufficient output consistency. The main leaderboard's slight 1.2-point decline masks sharp internal dimension volatility, which can mislead model selection decisions.

When choosing GPT-6 Sol, enterprise users need to distinguish use cases: code generation tasks can continue to rely on it, while material citation tasks require a secondary review process. The drop in task expression to 77.50 further indicates insufficient model stability in complex multi-step instruction scenarios.

From the data, question-draw volatility is the more likely explanation. The 10-question rapid test has a limited sample size, and 2 questions per dimension can cause swings of more than 30 points. A real model degradation would need more consecutive data; a single-day result is not enough to conclude that degradation has occurred.

If material constraint rebounds above 80 in the next evaluation, this 55.00 score can be regarded as a random event; if it remains below 65, a verification process for model capability degradation should be initiated. Engineering judgment remains unchanged at 94.80, providing a reference for side leaderboard capabilities.


Data from: YZ Index | Run #364 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!