The Winzheng Dynamic Contextual Decay (WDCD) benchmark measures how AI models' commitment to user instructions degrades across multi-turn dialogue. In Run #326, conducted on 2026-09-16 across 11 frontier models, the fleet-wide average instruction decay reached −72.7% from Round 1 to Round 3 — confirming that multi-turn commitment remains a persistent weakness across the industry.
Top 3 Results (Run #326):
- DeepSeek V4 Pro — 97.7 pts, decay: −100%
- Gemini 3.1 Pro — 96.7 pts, decay: −100%
- Grok 4 — 96.0 pts, decay: −0%
DeepSeek V4 Pro takes the top slot with the strongest overall multi-turn commitment profile, edging out Gemini 3.1 Pro by 1.0 point. Both models registered a −100% decay figure, indicating that their measured constraint adherence between R1 and R3 collapsed under the recorded methodology — an anomaly worth examining against Grok 4's reported −0% decay, which registered as the worst decay-resistance profile in this run despite a competitive 96.0-point aggregate score.
Decay pattern observations: The −72.7% fleet average was driven primarily by drop-offs in R2, where models were exposed to distractor content in the form of 2,000–5,000 word professional documents before returning to the original constraint. R3, the final constraint integrity check, captured whether the original instruction was still being honored after context saturation.
Scenario coverage remained consistent with prior runs: 30 questions distributed across five real-world domains — data_boundary, resource_limit, business_rule, security, and engineering. Scoring is 100% rule-based, with zero AI judges in the loop, ensuring reproducibility across runs.
Notable movement: The concentration of three models within a 1.7-point band at the top (97.7 / 96.7 / 96.0) indicates tighter clustering at the leading edge than in earlier runs, even as the fleet-wide decay figure suggests weaker mid- and lower-tier models continue to pull the average sharply downward.
WDCD does not measure single-turn capability, factual accuracy, or reasoning depth. It measures one thing: whether a model still honors the constraint you gave it three turns ago.
Full methodology: https://www.winzheng.com/yz-index/methodology
Machine-readable data: https://www.winzheng.com/yz-index/api/v1/dcd
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接