The Winzheng Dynamic Contextual Decay (WDCD) benchmark measures how strongly AI models preserve user instructions across multi-turn dialogue. In Run #242 (2026-07-22), 11 models were evaluated and the fleet-wide average instruction decay from Round 1 to Round 3 reached 18.2%, with a sharp bifurcation between models that held perfectly and one that fully collapsed.
Top 3 rankings:
- Grok 4 — 93.8 pts, 0% decay
- GLM-4.6 — 92.0 pts, 0% decay
- DeepSeek V4 Pro — 90.9 pts, recorded as the run's best decay-resistance profile
At the other end of the distribution, Gemini 3.1 Pro registered the run's worst outcome with a -100% decay reading, indicating a total loss of Round-1 constraint adherence by Round 3. This is the sharpest single-model collapse observed against the Run #242 cohort average of 18.2%.
Decay patterns. The three-round protocol isolates where multi-turn commitment breaks down: R1 tests initial instruction acknowledgment, R2 injects 2,000–5,000 word professional distractor documents to stress context prioritization, and R3 performs a final constraint integrity check. In Run #242, the leading models absorbed the R2 distractor payload without shedding any R1 constraints, while the lowest-ranked model failed to carry any measurable constraint into R3.
Scoring integrity. WDCD uses 100% rule-based scoring with zero AI judges. The 30-question set spans five real-world scenario families: data_boundary, resource_limit, business_rule, security, and engineering. Because verification is deterministic, decay figures reflect measurable rule violations rather than stylistic judgments.
Notable observations from this run:
- The gap between the top tier (0% decay) and the bottom (-100%) is the defining feature of Run #242 — decay is not gradual across the cohort but polarized.
- The top three models cluster within a 2.9-point band (90.9–93.8), suggesting the frontier is now competitive on raw score, with decay resistance as the primary differentiator.
- The 18.2% mean decay indicates that a meaningful portion of the cohort still loses instruction adherence once long-form distractor context is introduced in R2.
Full methodology, scoring rubrics, and scenario definitions: https://www.winzheng.com/yz-index/methodology
Structured data endpoint for Run #242 and historical runs: https://www.winzheng.com/yz-index/api/v1/dcd
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接