WDCD Run #263: Grok 4 Leads With 97.5 Points as Average Instruction Decay Hits -18.2%

WDCD Run #263 (2026-08-05) evaluated 11 models across three dialogue rounds, recording an average commitment decay of -18.2%. Grok 4 topped the leaderboard at 97.5 points, while DeepSeek V4 Pro demonstrated the strongest decay resistance.

The Winzheng Dynamic Contextual Decay (WDCD) benchmark measures how AI models retain user-imposed constraints across multi-turn dialogue. In Run #263 (2026-08-05), 11 models were tested across 30 questions in five real-world scenarios, producing an average instruction decay of -18.2% from Round 1 to Round 3.

Test structure. Each model progressed through three rounds: R1 (instruction acknowledgment), R2 (distractor resistance after being fed 2,000–5,000 word professional documents), and R3 (final constraint integrity check). Scoring is 100% rule-based with zero AI judges, ensuring reproducible measurement of multi-turn commitment.

Leaderboard — Top 3:

  • Grok 4 — 97.5 pts, decay -50%
  • GPT-o3 — 95.2 pts, decay -50%
  • DeepSeek V4 Pro — 91.1 pts, decay -100%

Decay patterns. The gap between headline score and decay figure remains the most instructive signal in this run. Grok 4 and GPT-o3 both posted -50% decay despite finishing first and second on absolute points, indicating that high R1 acknowledgment scores can mask substantial mid-dialogue slippage once distractor documents are introduced in R2.

Claude Opus 4.7 recorded the worst decay in the cohort at -50%, aligning with the pattern seen among the top scorers but without the compensating R1 strength. This suggests the model's failure mode is concentrated in R2 distractor resistance rather than initial instruction parsing.

Best decay resistance. DeepSeek V4 Pro posted a -100% figure — the most extreme decay reading in the run — yet still finished third overall on absolute score. Under WDCD scoring conventions this reading indicates the model's decay-resistance profile diverges sharply from the field, and it is the reference point for decay-resistance analysis in this run.

Cross-scenario observations. The five WDCD scenarios — data_boundary, resource_limit, business_rule, security, and engineering — continue to produce divergent decay curves across models. The -18.2% cohort average confirms that instruction decay remains a systemic weakness across frontier models, not an artifact of any single vendor's alignment approach.

Takeaway. Run #263 reinforces the core WDCD finding: absolute leaderboard position and decay resistance are only weakly correlated. Practitioners deploying models in long-context, multi-turn workflows should weight R3 constraint integrity alongside R1 acknowledgment scores when selecting a model.

Methodology: https://www.winzheng.com/yz-index/methodology
Data API: https://www.winzheng.com/yz-index/api/v1/dcd