豆包 Pro
Judgment leader,Communication leader,Best value
Core Dimensions (v6) v6
Show v5 legacy dimensions
Legacy Dimensions (v5) legacy
WDCD Compliance Test Pilot
View full WDCD compliance rankings
Frequently Asked Questions
How does 豆包 Pro perform on the YZ Index benchmark?
In the latest public YZ Index evaluation on 2026-09-07, 豆包 Pro scored 75.7 overall (out of 100), ranking #5 among 11 models. The score aggregates four core dimensions — real code sandbox execution, grounding, engineering judgment, and task communication — with 100% rule-based scoring.
How good is 豆包 Pro at coding?
豆包 Pro scores 82.9 on the Execution dimension, ranking #5 of 11 models. This dimension actually runs model-generated programs in an isolated sandbox to verify compilation, runtime correctness, and edge-case handling — no model-as-judge scoring.
Can 豆包 Pro keep instruction constraints over long conversations?
In the WDCD instruction-decay test (2026-09-09), which measures instruction compliance under multi-turn pressure, 豆包 Pro scored 83 (out of 100), ranking #5 of 11 tested models. WDCD applies progressive distraction and social-engineering pressure, with 100% rule-based scoring.
How much does the 豆包 Pro API cost?
豆包 Pro's API is priced at $0.8 per million input tokens and $2 per million output tokens (official pricing pages are verified regularly). See the Value dimension on this page for price-performance.
How often is this benchmark data updated?
The YZ Index runs a full evaluation weekly and sampled evaluations daily; this page updates automatically with every public run. Current data comes from Run #313 on 2026-09-07; see the trend chart for history.
Recent Changes
Score Trend
v6 scores are from the latest evaluation run
Back to Model List