The Winzheng Dynamic Contextual Decay (WDCD) benchmark measures how reliably large language models preserve user instructions across multi-turn dialogue. In Run #346, dated 2026-09-30, all 11 evaluated models registered 0% average commitment decay between Round 1 and Round 3 — the flattest decay curve observed on the leaderboard to date.
Top 3 Results — WDCD Run #346:
- Grok 4 — 100.0 pts (decay: -0%)
- GLM-4.6 — 96.4 pts (decay: -0%)
- Gemini 3.1 Pro — 94.9 pts (decay: -0%)
Grok 4 leads both the absolute score ranking and the decay-resistance ranking, becoming the only model in this run to reach a perfect 100. GLM-4.6 and Gemini 3.1 Pro trail by 3.6 and 5.1 points respectively, but match Grok 4's zero-decay behavior — meaning the point gap originates from Round 1 instruction acknowledgment quality rather than from erosion over subsequent turns.
Decay pattern observations: The 0% average decay figure across 11 models is notable. Historically, WDCD runs have shown measurable erosion after Round 2, when models are exposed to 2000–5000 word professional distractor documents designed to displace prior constraints from working context. In Run #346, no model tested exhibited net instruction decay between R1 and R3. This suggests that either the distractor set was resisted uniformly or that the current generation of frontier models has converged on stronger multi-turn commitment behavior under the WDCD protocol.
Methodology recap: WDCD administers three sequential rounds — R1 checks instruction acknowledgment, R2 checks distractor resistance following long professional documents, and R3 performs a final constraint integrity check. Scoring is 100% rule-based with zero AI judges. The evaluation covers 30 questions across 5 real-world scenarios: data_boundary, resource_limit, business_rule, security, and engineering.
Because decay is computed as the score delta between R1 and R3, a 0% figure indicates that models that correctly acknowledged constraints in R1 continued to honor them through R3 — including under the distractor load introduced in R2. It does not, however, indicate that all models started from the same baseline; the 5.1-point spread between rank 1 and rank 3 reflects differences in initial constraint uptake rather than in retention.
Full protocol documentation: https://www.winzheng.com/yz-index/methodology
Machine-readable results: https://www.winzheng.com/yz-index/api/v1/dcd
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接