DeepSeek V4 Pro Code Execution Drops 25 Points, Main Benchmark Slides 6.7 Points

DeepSeek V4 Pro’s code execution score dropped from 100.00 to 75.00 in today’s Smoke evaluation, while its main benchmark score fell from 83.53 to 76.85.

Score Facts and Dimensional Breakdown

The main benchmark is a weighted average of code execution and material constraint. Today’s code execution scored 75.00 and material constraint 79.10, yielding a main benchmark of 76.85. Code execution fell 25 points in a single day, while material constraint rose 15.7 points. Engineering judgment improved from 70.80 to 100.00, task expression dropped from 91.70 to 74.20, and integrity rating shifted from warn to pass.

Analysis of Fluctuation Causes

The Smoke evaluation uses only 10 questions per day (2 per dimension), so the small sample size naturally leads to higher daily standard deviation. The sharp drop in code execution is most likely due to the two questions selected today requiring higher demands in function calling, boundary conditions, or multi-step reasoning, while yesterday’s questions were relatively easier. The rise in material constraint suggests the model improved its faithfulness to given materials in today’s questions. Engineering judgment and task expression, both part of the side benchmark (AI-assisted evaluation), also saw drastic changes driven by the small sample, and cannot be directly interpreted as a structural shift in model capability.

Code execution: 75.00 vs. yesterday’s 100.00; Material constraint: 79.10 vs. yesterday’s 63.40; Main benchmark: 76.85 vs. yesterday’s 83.53.

Current data does not support a conclusion of “genuine model degradation,” because material constraint and engineering judgment both moved significantly upward, indicating that the overall output quality of the model did not decline systematically. More likely, the variance originates from question sampling.

Practical Implications for Users

Teams that heavily depend on code execution (e.g., developers automating test case generation or data processing scripts) should add manual review steps after today’s results to avoid potential errors introduced by daily fluctuations. A material constraint score of 79.10 suggests the model performs adequately in scenarios that require strict adherence to given documents or API specifications, and can continue to be used for document-based Q&A tasks. The integrity rating turning to pass lowers the compliance barrier for deploying the model in production environments.

Strategic Assessment

Based on the current score comparison, the main benchmark decline of DeepSeek V4 Pro is primarily driven by variance in the single dimension of code execution, not a full-dimensional degradation. If code execution remains below 80 in the next evaluation, attention priority should be raised; if it rebounds above 90, today’s result can be judged as typical sampling noise. The current data only supports a “continue to monitor” conclusion, not a downgrade in model priority.


Data source: YZ Index | Run #254 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!