DeepSeek V4 Pro Scores 55.60 on Material Constraint, Code Execution Dimension Missing, Misses Main Ranking in Smoke Test

DeepSeek V4 Pro had missing data in the code execution dimension in today's Smoke evaluation, scoring 55.60 on material constraint, and was therefore unable to participate in the main ranking.

Data Facts: Dimension Scores and Missing Status

Today's score comparison shows no data at all for the code execution dimension, a material constraint score of 55.60, engineering judgment (side leaderboard, AI-assisted evaluation) at 50.00, and task expression (side leaderboard, AI-assisted evaluation) at 62.50. No scores were recorded for any dimension yesterday, and the main leaderboard for this cycle is marked as missing overall.

Cause Analysis: Distinguishing API Failure from Question Draw Fluctuation

The material constraint score of 55.60 and engineering judgment score of 50.00 point to API calls being interrupted during the execution phase. The code execution dimension requires the model to run code directly and return results; an API timeout would directly leave this dimension empty, while material constraint only requires text-level alignment, allowing a score of 55.60 to still be produced. Task expression at 62.50 is higher than engineering judgment, indicating that the model is relatively stable on instruction-restatement tasks, but engineering judgment involves multi-step logic chains, and the 50.00 score shows insufficient consistency in the side leaderboard assessment.

The Smoke evaluation uses only 2 questions per dimension per day, and the question draw itself carries inherent randomness. If the model had genuinely degraded, material constraint and task expression would typically decline in tandem; this time, however, only the execution dimension is missing while other dimensions retain values, pointing to an API failure rather than an overall decline in model capability. The stability dimension measures the standard deviation of scores; since the execution dimension was not measured this cycle, the standard deviation cannot be calculated, making it impossible to determine its consistency level.

Implications for Users

Teams with heavy code execution workloads should add a timeout retry mechanism when calling DeepSeek V4 Pro. The material constraint score of 55.60 indicates it remains usable in document-fidelity scenarios, but the missing execution dimension will directly impact code-generation applications. Developers relying on this model should prioritize verifying API availability before assessing the impact of the 50.00 engineering judgment score on complex tasks.

For scenarios sensitive to material fidelity, 55.60 provides a quantifiable baseline; if the task also requires code execution verification, the data from this cycle is insufficient to support a model selection decision.

Strategic Assessment

Based on the available score comparison, the anomaly for DeepSeek V4 Pro this cycle is primarily caused by an API failure rather than genuine model degradation. The failure to secure a main leaderboard ranking warrants continued verification in the next cycle as to whether the execution dimension recovers. The gap between the engineering judgment score of 50.00 and task expression score of 62.50 suggests an internal imbalance in side leaderboard capabilities that requires further observation in subsequent complete data.

No anomaly appeared in the integrity rating in this cycle's data, so it does not currently affect the entry threshold assessment. As operational signals, stability and usability cannot be assigned specific scores this cycle due to incomplete data.


Data Source: YZ Index | Run #279 | View Raw Data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!