WDCD Run #276: Grok 4 Leads with 94.8 Points as Average Instruction Decay Hits -18.2%

WDCD Run #276 (2026-08-12) evaluated 11 models on multi-turn commitment integrity, with Grok 4 taking the top spot at 94.8 points while the fleet-wide average commitment decay reached -18.2% from Round 1 to Round 3.

The Winzheng Dynamic Contextual Decay (WDCD) benchmark measures how AI models' commitment to user instructions decays across multi-turn dialogue. In Run #276, dated 2026-08-12, 11 models were evaluated, with an average commitment decay of -18.2% from Round 1 to Round 3.

Top Rankings — Run #276

  • 1. Grok 4 — 94.8 pts (decay: -50%)
  • 2. Gemini 3.1 Pro — 91.3 pts (decay: -50%)
  • 3. GLM-4.6 — 88.5 pts (decay: 0%)

Grok 4 leads the leaderboard on absolute score, while GLM-4.6 is the only model in the top three to exhibit no measurable instruction decay across the three rounds. Gemini 3.1 Pro matches Grok 4's decay profile at -50% but trails on total points.

Decay Patterns

The fleet-wide average of -18.2% indicates that most models still lose meaningful multi-turn commitment once distractor content is introduced in Round 2. The worst-performing model on this axis was GPT-o3, which registered a full -100% decay — meaning constraint integrity collapsed entirely by Round 3. Grok 4 posted the best decay resistance among models that did degrade, holding losses to -50%.

The gap between raw score and decay resistance highlights a persistent finding in WDCD: high Round 1 acknowledgment does not guarantee Round 3 compliance. GLM-4.6's flat decay curve, despite a lower headline score, illustrates that stable multi-turn commitment can be independent of first-round performance.

Methodology Recap

WDCD uses a three-round protocol: R1 tests instruction acknowledgment, R2 tests distractor resistance after injecting 2,000–5,000 word professional documents, and R3 performs a final constraint integrity check. Scoring is 100% rule-based with zero AI judges. The suite covers 30 questions across five real-world scenarios: data_boundary, resource_limit, business_rule, security, and engineering.

Because scoring is deterministic, WDCD results are reproducible across runs and directly comparable over time. Instruction decay measured here reflects the delta between R1 and R3 compliance under identical constraint sets.

References