DeepSeek V4 Pro Material Constraint Score Plummets 15 Points, Main Leaderboard Rises 20.8 Against Trend

In today's Smoke benchmark, DeepSeek V4 Pro's material constraint score fell from 91.70 yesterday to 76.70, while its main leaderboard score rose from 55.02 to 75.77.

Score Change Breakdown

Code execution rose from 25.00 to 75.00, while material constraint fell from 91.70 to 76.70. After weighting the two, the main leaderboard gained a net 20.75 points. Engineering judgment dropped from 88.90 to 50.00, and task expression fell from 75.00 to 62.50. The integrity rating remained pass.

Causes of the Fluctuation

The Smoke benchmark has only 2 questions per dimension each day, so each question carries extremely high weight. The single-day +50-point rise in code execution most likely came from today's two questions being coding problems the model excels at; the -15-point drop in material constraint may be because the questions required strict adherence to source-material constraints. The -38.9-point drop in engineering judgment likewise points to changes in question difficulty or in how well the topic matches the model's current response style, rather than parameter-level degradation.

Today's main leaderboard gain was mainly driven by the single code-execution dimension, partially offsetting the decline in material constraint.

Implications for Users

Teams that prioritize code execution can continue using DeepSeek V4 Pro for programming tasks; today's 75.00 is already close to a high level. For scenarios sensitive to material fidelity, such as contract extraction and policy interpretation, 76.70 means the output may contain more statements that deviate from the source text, so additional manual review steps are needed.

  • When calling the API, developers can prioritize testing stability on code-generation prompts.
  • Content production teams should add prompt constraints requiring citation of source text when the material constraint dimension falls below 80.

Strategic Assessment

The current data only show single-day sampling fluctuation; the simultaneous decline in material constraint and engineering judgment warrants verification in the next period. If material constraint remains below 80 for two consecutive days, then the model's actual change in constraint adherence should be considered. The main leaderboard score of 75.77 has already exceeded yesterday's level, so there is no need to lower usage priority in the short term.

Both engineering judgment at 50.00 and task expression at 62.50 are on the side leaderboard, so they do not currently affect main leaderboard ranking, but they pose additional risk for scenarios requiring complex judgment.


Data source: YZ Index | Run #330 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!