DeepSeek V4 Pro Code Execution Plunges by 23.7 Points, Main Leaderboard Drops 5.2 Points

DeepSeek V4 Pro in today's Smoke Benchmark saw its Code Execution dimension fall from 98.70 to 75.00 points, a drop of 23.7 points, and the overall main leaderboard declined from 91.46 to 86.25 points.

Score Comparison Data

The Material Constraint dimension rose from 82.60 to 100.00 points, an increase of 17.4 points. Engineering Judgment dropped from 100.00 to 88.90 points, a decline of 11.1 points. Task Expression rose from 81.70 to 91.70 points, an increase of 10 points. Integrity Rating remained at pass.

Causal Analysis: Sampling Fluctuation or Real Degradation

The Smoke Benchmark uses only 10 questions per day, with 2 questions per dimension, leading to naturally large daily standard deviations. The concentrated score loss in the Code Execution dimension this time may stem from drawing questions requiring complex multi-step debugging or edge cases, whereas yesterday’s questions were relatively simple. The Material Constraint dimension reversely hit a full score, indicating stable model performance on citation adherence and fidelity. Two dimensions showing extreme opposite changes simultaneously is more consistent with random question-sampling fluctuation than systematic degradation of model parameters or training.

The Engineering Judgment dimension dropped 11.1 points in the same direction as Code Execution, while Task Expression rose 10 points, partially offsetting each other to yield a net 5.2-point drop on the main leaderboard. The overall fluctuation range remains within normal bounds for daily quick tests.

Implications for Users

Enterprises heavily dependent on code execution scenarios—such as automated script generation or algorithm prototype validation teams—should increase manual review steps today, avoiding direct adoption of outputs under low scores. The Material Constraint dimension reaching a full score provides higher confidence for document generation scenarios that require strict source citation or hallucination avoidance.

With Engineering Judgment dropping to 88.90 points, projects requiring architectural decisions or multi-scheme trade-offs should consider temporarily switching to more stable models for cross-validation.

Strategic Assessment

Based on single-day comparison, the code execution plunge is more likely caused by question-sampling fluctuation than real model degradation. Only if similar declines occur for two or more consecutive days on the same dimension should the stability of this Smoke Benchmark run be closely scrutinized. Current data does not support a long-term downgrade of DeepSeek V4 Pro's code capabilities.


Data Source: YZ Index (YZ Index) | Run #232 | View Raw Data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!