DeepSeek V4 Pro saw its code execution dimension fall from 100.00 to 75.00 points in today's Smoke evaluation, while the material constraint dimension rose from 68.20 to 95.00 points. The main leaderboard score changed from 85.69 to 84.00 points.
Data Facts: Dimensional Scores and Leaderboard Changes
The code execution dimension dropped 25 points in a single day, while the engineering judgment dimension fell from 94.50 to 75.00 points. The task expression dimension slightly declined by 4.7 points to 90.00. The material constraint dimension rose by 26.8 points, directly offsetting the loss in code execution, resulting in a main leaderboard decline of only 1.7 points. The integrity rating remained at pass.
Cause Analysis: Question Sampling Fluctuation or Model Degradation
The Smoke evaluation tests only 10 questions per day, with 2 questions per dimension. The sample size is extremely small, so the standard deviation of daily scores is naturally large. The simultaneous drop of over 19.5 points in both code execution and engineering judgment most likely indicates that the two code questions sampled that day were significantly harder than the previous day, causing the model to make errors in boundary condition handling or multi-step reasoning chains. The sharp rise in the material constraint dimension suggests that the material questions that day were closer to the model's training distribution, reducing hallucinations or formatting violations.
If this were a genuine model degradation, it would typically be accompanied by simultaneous and sustained declines across multiple dimensions, rather than an opposing sharp rise in material constraints. The current data only shows a single day of contrasting changes, with no evidence of consecutive days of weakening in the same dimension. Therefore, it is more likely a question sampling fluctuation.
Implications for Users
Teams that heavily rely on code execution should add local testing steps when calling DeepSeek V4 Pro, especially when handling complex algorithms or multi-file interaction scenarios. A single 75-point performance may lead to a lower pass rate for generated code. The material constraint score of 95 indicates that this model is more robust when strictly following user-provided documents or format requirements, making it suitable for document generation or data extraction tasks that demand high-fidelity output.
The engineering judgment dimension dropping to 75 points suggests that developers who need the model to assist with architectural decisions or code reviews should reduce trust in single outputs and increase manual review steps.
Strategic Judgment
The main leaderboard only dropped 1.7 points, masking the real volatility in code execution and engineering judgment, indicating that the current leaderboard formula gives a higher weight to the material constraint dimension. DeepSeek V4 Pro's single-day code execution performance has returned to the 75-point range. If this dimension fails to recover to above 90 points in the next evaluation, it will be necessary to verify whether a capability degradation signal has emerged. Based solely on a single day of data, it is more reasonable to classify this as normal sampling fluctuation.
The opposing trends of code execution at 75.00 points and material constraint at 95.00 points are the most noteworthy phenomenon in this evaluation.
For enterprise selection, DeepSeek V4 Pro can still serve as a candidate model for material constraint-priority scenarios, but backup options should be retained for code generation pipelines. If both code execution and engineering judgment recover in the next Smoke evaluation, this round can be confirmed as random fluctuation; if they remain low, further observation is needed.
Data source: YZ Index | Run #250 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接