DeepSeek V4 Pro in today's Smoke evaluation saw its Material Constraint dimension score drop from 100.00 yesterday to 68.20 points, a decline of 31.8 points; its Code Execution dimension rose from 69.50 points to 100.00 points, an increase of 30.5 points.
Data Facts: Minor Rise in Main Scoreboard Conceals Severe Dimensional Shifts
The Main Scoreboard score rose from 83.23 to 85.69 points, a net increase of 2.46 points. Engineering Judgment dropped from 100.00 to 94.50 points, and Task Expression fell from 100.00 to 94.70 points. Integrity Rating remained at pass.
The Smoke evaluation uses only 2 questions per dimension per day, resulting in an extremely small daily sample size. The near-symmetric 30-point reverse movements in Material Constraint and Code Execution point to question assignment randomness rather than a structural degradation in model capability.
Cause Analysis: Sampling Fluctuation, Not Genuine Degradation
Material Constraint scored full marks yesterday but 68.20 today, a difference of 31.8 points; Code Execution scored 69.50 yesterday but 100.00 today, a difference of 30.5 points. The magnitude of change for the two dimensions is similar and opposite in direction. The most likely explanation is that the two Material Constraint questions drawn today had significantly higher difficulty or stricter constraints than yesterday's, while the Code Execution questions were relatively easier.
Engineering Judgment and Task Expression each dropped by 5.5 and 5.3 points, much smaller than the main dimensions, consistent with small-sample noise characteristics. If the model had shown a systemic degradation in material fidelity, it would typically be accompanied by a simultaneous decline in Code Execution, not a 30.5-point offsetting increase.
Implications for Users
For enterprises or developers heavily reliant on material citations, today's 68.20 points mean a notably higher probability of single-call failure in scenarios requiring strict fidelity to the original text. Additional manual review steps or multi-round verification are needed.
Teams focused on code execution can enjoy greater certainty. Today's 100.00 points show the model has reached a perfect level in algorithm implementation and debugging tasks, making it a priority for code generation pipelines.
Strategic Assessment
A single-day 31.8-point drop in Material Constraint falls within the normal fluctuation range of Smoke evaluations and does not constitute grounds for an immediate rating downgrade. The Main Scoreboard still posted positive growth, indicating that the perfect Code Execution score has fully offset the Material Constraint loss.
If the Material Constraint score continues to stay in the sub-70 range in the next period, while Code Execution remains high, it will be necessary to verify whether there have been changes in the material processing mechanism at the prompt template or system instruction level. Current data only supports the conclusion of "caused by sampling," not the judgment of "model degradation."
Data Source: YZ Index | Run #248 | View Raw Data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接