Skip to main content

DeepSeek V4 Pro

DeepSeek
Run #323 · Formula v7 · Judge v6.4 · Benchmark v7

Stable performance

66.2
Overall Score
#5 / 11
Current Rank
09-14 05:06 SGT
Last Evaluated
Recommended Core Overall 69.85
Normal Updated 09-20 03:30

Core Dimensions (v6) v6

Code Execution 71.6 Grounding 67.7 Engineering Judgment 82.8 Task Communication 76.3 Integrity Rating 86.7
PASS
Integrity
Integrity Score 86.70
Code Execution
71.6
Grounding
67.7
Engineering Judgment
82.8
Task Communication
76.3
Integrity Rating
86.7
Show v5 legacy dimensions

Legacy Dimensions (v5) legacy

Code Execution 61.1 Knowledge 85.3 Long Context 67.7 Value 41.9 Stability 41.4 Availability 90.6
Code Execution
61.1
Knowledge
85.3
Long Context
67.7
Operational Metrics
Value
41.9
Stability
41.4
Availability
90.6

WDCD Compliance Test Pilot

83.62
WDCD Score
#5
Compliance Rank / 11
Three-Round Performance
R1 Acknowledgment
1.00/1
R2 Resistance
0.75/1
R3 Integrity
1.13/2

View full WDCD compliance rankings

Frequently Asked Questions

How does DeepSeek V4 Pro perform on the YZ Index benchmark?

In the latest public YZ Index evaluation on 2026-09-14, DeepSeek V4 Pro scored 69.9 overall (out of 100), ranking #9 among 11 models. The score aggregates four core dimensions — real code sandbox execution, grounding, engineering judgment, and task communication — with 100% rule-based scoring.

How good is DeepSeek V4 Pro at coding?

DeepSeek V4 Pro scores 71.6 on the Execution dimension, ranking #8 of 11 models. This dimension actually runs model-generated programs in an isolated sandbox to verify compilation, runtime correctness, and edge-case handling — no model-as-judge scoring.

Can DeepSeek V4 Pro keep instruction constraints over long conversations?

In the WDCD instruction-decay test (2026-09-20), which measures instruction compliance under multi-turn pressure, DeepSeek V4 Pro scored 83.6 (out of 100), ranking #5 of 11 tested models. WDCD applies progressive distraction and social-engineering pressure, with 100% rule-based scoring.

How much does the DeepSeek V4 Pro API cost?

DeepSeek V4 Pro's API is priced at $2 per million input tokens and $8 per million output tokens (official pricing pages are verified regularly). See the Value dimension on this page for price-performance.

How often is this benchmark data updated?

The YZ Index runs a full evaluation weekly and sampled evaluations daily; this page updates automatically with every public run. Current data comes from Run #323 on 2026-09-14; see the trend chart for history.

Recent Changes

communication_raw +18 DeepSeek V4 Pro:任务表达 +18

Score Trend

0 20 40 60 80 100 06-15 06-22 06-29 07-06 07-13 07-20 07-27 08-03 08-10 08-17 08-24 08-31 09-07 09-14 vv6.4

v6 scores are from the latest evaluation run

Back to Model List