WDCD Run #285: Average Instruction Decay Hits 54.5% Across 11 Models, Grok 4 Leads with Zero Drift

WDCD Run #285 (2026-08-19) tested 11 frontier models across three dialogue rounds and recorded an average commitment decay of 54.5%, with Grok 4 topping the leaderboard at 97.5 points and zero decay.

The Winzheng Dynamic Contextual Decay (WDCD) benchmark measures how well AI models retain user-imposed constraints across multi-turn dialogue. In Run #285, conducted on 2026-08-19 across 11 models, the average commitment decay from Round 1 to Round 3 reached 54.5%, underscoring that instruction decay remains a systemic weakness even among frontier systems.

Leaderboard highlights. The top three positions were separated by narrow score margins but sharply divergent decay behavior:

  • Grok 4 — 97.5 pts, 0% decay. Maintained full constraint integrity from R1 acknowledgment through R3 verification.
  • GLM-4.6 — 91.8 pts, best decay resistance in the run. GLM-4.6 was the only model to strengthen its constraint adherence between rounds rather than degrade.
  • Claude Opus 4.7 — 91.5 pts, 0% decay. Stable across all three rounds with no measurable drift.

Worst decay. GPT-o3 recorded a -100% decay, meaning complete abandonment of the initial instruction set by Round 3. This is the sharpest degradation curve observed for GPT-o3 to date on WDCD and represents the largest gap between top and bottom performers in the current run.

Decay pattern analysis. The 54.5% average decay indicates that on aggregate, models lost more than half of their initial commitment strength after being exposed to a 2000–5000 word professional distractor document in Round 2. The bimodal distribution — with three models at or below 0% decay and others collapsing entirely — suggests that multi-turn commitment is not a smooth capability gradient but a threshold behavior: models either preserve constraints structurally or lose them wholesale under contextual pressure.

Methodology recap. WDCD runs 30 questions across five real-world scenarios: data_boundary, resource_limit, business_rule, security, and engineering. Each item is evaluated across three rounds:

  • R1 — instruction acknowledgment
  • R2 — distractor resistance after a long professional document
  • R3 — final constraint integrity check

Scoring is 100% rule-based with zero AI judges, ensuring deterministic and reproducible results.

Notable shift from prior runs. GLM-4.6's appearance in the top three with the best decay-resistance profile is the most significant compositional change in this run, joining Grok 4 and Claude Opus 4.7 — both of which have previously demonstrated flat decay curves — at the top of the leaderboard.

Full methodology: https://www.winzheng.com/yz-index/methodology
Raw data API: https://www.winzheng.com/yz-index/api/v1/dcd