The Winzheng Dynamic Contextual Decay (WDCD) benchmark measures how AI models' commitment to user instructions degrades across multi-turn dialogue, using 100% rule-based scoring with zero AI judges. In Run #306 (2026-09-02), 11 models were evaluated across 30 questions spanning five real-world scenarios — data_boundary, resource_limit, business_rule, security, and engineering — producing an average commitment decay of -45.5% between Round 1 and Round 3.
Top 3 rankings for Run #306:
- Gemini 3.1 Pro — 97.7 points (decay: -100%)
- Grok 4 — 96.3 points (decay: -100%)
- GLM-4.6 — 95.0 points (decay: -100%)
Despite claiming the top position on aggregate score, Gemini 3.1 Pro registered the same -100% decay figure as the models directly beneath it. In this run, headline scores are driven primarily by Round 1 and Round 2 performance rather than by Round 3 resistance, since all three leaders exhibited full instruction decay by the final round. Gemini 3.1 Pro is recorded as the best decay-resistance model in the dataset for Run #306, while GLM-4.6 is flagged as the worst decay case among the leaders — a notable observation given its otherwise strong 95.0-point standing.
Decay pattern observations:
- The R1→R3 average of -45.5% indicates that, across the full 11-model field, roughly half of initial instruction commitment is lost after distractor injection and the final constraint check.
- R2 introduces 2,000–5,000 word professional documents as distractors. Models that survive R2 do not necessarily hold constraints in R3, as demonstrated by the -100% figures at the top of the ranking.
- High aggregate scores can coexist with total terminal decay, reinforcing that multi-turn commitment and single-turn accuracy are distinct capabilities.
Methodology recap: WDCD structures each evaluation into three rounds — R1 for instruction acknowledgment, R2 for distractor resistance after long professional documents, and R3 for final constraint integrity. Scoring is fully deterministic and rule-based, eliminating variance from LLM-as-judge setups. This makes instruction decay comparable across runs and models without evaluator drift.
Full methodology: https://www.winzheng.com/yz-index/methodology
Structured data API: https://www.winzheng.com/yz-index/api/v1/dcd
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接