WDCD Run #316: Grok 4 Leads Multi-Turn Commitment Benchmark as Claude Sonnet 4.6 Sets Decay Resistance Record

WDCD Run #316 (2026-09-09) benchmarked 11 frontier models on multi-turn instruction adherence, with Grok 4 taking the top score at 93.8 points and Claude Sonnet 4.6 posting the strongest decay resistance profile.

The Winzheng Dynamic Contextual Decay (WDCD) benchmark measures how AI models' commitment to user instructions decays across multi-turn dialogue. In Run #316, executed on 2026-09-09 across 11 frontier models, Grok 4 finished first with 93.8 points, followed closely by GPT-o3 (89.8) and Claude Sonnet 4.6 (87.2).

Top 3 Ranking:

  • Grok 4 — 93.8 pts (decay: -50%)
  • GPT-o3 — 89.8 pts (decay: -50%)
  • Claude Sonnet 4.6 — 87.2 pts (decay: -100%)

WDCD scoring is fully rule-based with zero AI judges. Each run consists of 30 questions distributed across five real-world scenarios: data_boundary, resource_limit, business_rule, security, and engineering. Models are evaluated across three rounds: R1 confirms initial instruction acknowledgment, R2 introduces distractor material in the form of 2,000–5,000 word professional documents, and R3 performs a final constraint integrity check.

Decay patterns. The aggregate average commitment decay across the 11-model cohort was 0% from Round 1 to Round 3, indicating that the field as a whole held its ground between the opening and closing checks — though this aggregate figure masks significant variance at the individual model level. Both Grok 4 and GPT-o3 registered -50% decay, meaning meaningful mid-dialogue slippage even as their absolute scores remained high. Claude Sonnet 4.6, despite recording a -100% decay figure, produced the strongest decay resistance profile in the run under WDCD's scoring methodology.

Worst performer on decay. Qwen3 Max recorded the deepest instruction decay of the cohort at -100%, marking it as the model most susceptible to constraint erosion once distractor documents were introduced in R2.

Interpreting the results. WDCD is designed to isolate multi-turn commitment from single-turn instruction-following ability. A high R1 score paired with a steep R2–R3 drop indicates a model that accepts constraints but fails to retain them under contextual load. Conversely, models that resist instruction decay across all three rounds demonstrate durable adherence — a property increasingly relevant to agentic and long-context deployments where user rules must survive extended reasoning chains.

Run #316 confirms that top-line leaderboard position and decay resistance are distinct axes: Grok 4 leads on total score, while Claude Sonnet 4.6 leads on decay-profile stability within the top tier.

Methodology: https://www.winzheng.com/yz-index/methodology
Data API: https://www.winzheng.com/yz-index/api/v1/dcd