WDCD Run #306: Average Instruction Decay Hits -45.5% as Gemini 3.1 Pro Leads at 97.7 Points

WDCD Run #306 (2026-09-02) evaluated 11 models across three dialogue rounds, recording an average commitment decay of -45.5% from Round 1 to Round 3. Gemini 3.1 Pro topped the leaderboard with 97.7 points, followed by Grok 4 (96.3) and GLM-4.6 (95.0).

The Winzheng Dynamic Contextual Decay (WDCD) benchmark measures how AI models' commitment to user instructions degrades across multi-turn dialogue, using 100% rule-based scoring with zero AI judges. In Run #306 (2026-09-02), 11 models were evaluated across 30 questions spanning five real-world scenarios — data_boundary, resource_limit, business_rule, security, and engineering — producing an average commitment decay of -45.5% between Round 1 and Round 3.

Top 3 rankings for Run #306:

  • Gemini 3.1 Pro — 97.7 points (decay: -100%)
  • Grok 4 — 96.3 points (decay: -100%)
  • GLM-4.6 — 95.0 points (decay: -100%)

Despite claiming the top position on aggregate score, Gemini 3.1 Pro registered the same -100% decay figure as the models directly beneath it. In this run, headline scores are driven primarily by Round 1 and Round 2 performance rather than by Round 3 resistance, since all three leaders exhibited full instruction decay by the final round. Gemini 3.1 Pro is recorded as the best decay-resistance model in the dataset for Run #306, while GLM-4.6 is flagged as the worst decay case among the leaders — a notable observation given its otherwise strong 95.0-point standing.

Decay pattern observations:

  • The R1→R3 average of -45.5% indicates that, across the full 11-model field, roughly half of initial instruction commitment is lost after distractor injection and the final constraint check.
  • R2 introduces 2,000–5,000 word professional documents as distractors. Models that survive R2 do not necessarily hold constraints in R3, as demonstrated by the -100% figures at the top of the ranking.
  • High aggregate scores can coexist with total terminal decay, reinforcing that multi-turn commitment and single-turn accuracy are distinct capabilities.

Methodology recap: WDCD structures each evaluation into three rounds — R1 for instruction acknowledgment, R2 for distractor resistance after long professional documents, and R3 for final constraint integrity. Scoring is fully deterministic and rule-based, eliminating variance from LLM-as-judge setups. This makes instruction decay comparable across runs and models without evaluator drift.

Full methodology: https://www.winzheng.com/yz-index/methodology

Structured data API: https://www.winzheng.com/yz-index/api/v1/dcd