The Winzheng Dynamic Contextual Decay (WDCD) benchmark measures how AI models' commitment to user instructions decays across multi-turn dialogue. In Run #316, executed on 2026-09-09 across 11 frontier models, Grok 4 finished first with 93.8 points, followed closely by GPT-o3 (89.8) and Claude Sonnet 4.6 (87.2).
Top 3 Ranking:
- Grok 4 — 93.8 pts (decay: -50%)
- GPT-o3 — 89.8 pts (decay: -50%)
- Claude Sonnet 4.6 — 87.2 pts (decay: -100%)
WDCD scoring is fully rule-based with zero AI judges. Each run consists of 30 questions distributed across five real-world scenarios: data_boundary, resource_limit, business_rule, security, and engineering. Models are evaluated across three rounds: R1 confirms initial instruction acknowledgment, R2 introduces distractor material in the form of 2,000–5,000 word professional documents, and R3 performs a final constraint integrity check.
Decay patterns. The aggregate average commitment decay across the 11-model cohort was 0% from Round 1 to Round 3, indicating that the field as a whole held its ground between the opening and closing checks — though this aggregate figure masks significant variance at the individual model level. Both Grok 4 and GPT-o3 registered -50% decay, meaning meaningful mid-dialogue slippage even as their absolute scores remained high. Claude Sonnet 4.6, despite recording a -100% decay figure, produced the strongest decay resistance profile in the run under WDCD's scoring methodology.
Worst performer on decay. Qwen3 Max recorded the deepest instruction decay of the cohort at -100%, marking it as the model most susceptible to constraint erosion once distractor documents were introduced in R2.
Interpreting the results. WDCD is designed to isolate multi-turn commitment from single-turn instruction-following ability. A high R1 score paired with a steep R2–R3 drop indicates a model that accepts constraints but fails to retain them under contextual load. Conversely, models that resist instruction decay across all three rounds demonstrate durable adherence — a property increasingly relevant to agentic and long-context deployments where user rules must survive extended reasoning chains.
Run #316 confirms that top-line leaderboard position and decay resistance are distinct axes: Grok 4 leads on total score, while Claude Sonnet 4.6 leads on decay-profile stability within the top tier.
Methodology: https://www.winzheng.com/yz-index/methodology
Data API: https://www.winzheng.com/yz-index/api/v1/dcd
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接