DeepSeek V4 Pro Material Constraint Plummets 18.4 Points, Code Execution Rebounds 11.1 Points

In today's Smoke evaluation, DeepSeek V4 Pro's material constraint score plunged directly from 81.70 to 63.30, a decline of 18.4 points, while the main leaderboard overall dropped only 2.2 points, from 82.64 to 80.46.

Score Breakdown and Direct Evidence

The code execution dimension rose from 83.40 to 94.50, up 11.1 points; engineering judgment fell from 75.00 to 50.00, down 25 points; task expression declined from 75.00 to 66.70, down 8.3 points. The integrity rating remained "pass." All of the above data comes from the same daily 10-question Smoke rapid test, with 2 questions per dimension.

Analysis of Fluctuation Causes

The Smoke evaluation covers the material constraint dimension with only 2 questions per day; a single incorrect answer can lower that dimension by 20-30 points. The sharp drop in material constraint score most likely stems from the two questions drawn that day imposing higher demands on source fidelity, with the model exhibiting one clear instance of hallucination or excessive rewriting. The concurrent 11.1-point rise in code execution indicates that the model's overall reasoning capability has not undergone systematic degradation; this is more of a sampling fluctuation on specific constraint tasks.

The 25-point decline in engineering judgment likewise points to differences in question difficulty sampling rather than changes in model parameters. The main leaderboard's modest 2.2-point decline is precisely the result of the offsetting effects between code execution and material constraint, demonstrating that severe single-dimension fluctuations in small-sample rapid tests are a normal statistical phenomenon.

Specific Implications for Users

For teams heavily reliant on material constraint in RAG systems or enterprise knowledge bases, additional manual verification of DeepSeek V4 Pro outputs is needed today, particularly in scenarios involving quoting original text and avoiding fabricated details. The rise in code execution score indicates that the model performs more consistently on algorithm implementation and data processing script generation tasks.

For developers who also require engineering judgment, today's 50.00 score suggests reduced output consistency under architecture review or solution trade-off prompts. They should temporarily switch to other models or add multi-round confirmation steps.

Strategic Assessment

Based on the current score comparison, the 18.4-point decline in material constraint is most likely driven by question sampling fluctuation rather than genuine model degradation. The countervailing improvement in code execution further supports this assessment. Single-day data is insufficient to confirm a trend; it is recommended to continue observing in the next Smoke evaluation whether material constraint rebounds to the 80-point range. If it remains at a low level for two consecutive days, only then consider adjusting production environment invocation strategies.

At present, the main leaderboard score of 80.46 remains within a usable range. Lightweight tasks with lower stability requirements can continue using the model, while scenarios sensitive to material fidelity require added verification steps.


Data source: YZ Index | Run #292 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!