Instruction Compliance & WDCD

226 articles · Page 1 of 12
Does your AI model actually follow instructions? Instruction compliance is the most critical evaluation dimension for enterprise AI deployment, yet traditional benchmarks rarely test it. WDCD (Winzheng Dynamic Contextual Decay) is the world's first systematic test measuring how AI models' commitment to instructions decays over extended dialogue — using three rounds of 2,000-5,000 word professional distractions across 30 constraint questions in 5 real-world scenarios, with 100% rule-based scoring and zero AI judges. The YZ Index Integrity Rating also deploys 42 canary probes to detect fabricated citations and hallucinated data. This topic covers instruction compliance research, hallucination detection methods, and WDCD test result analysis.

In-depth Guides

Lab WDCD Run #360: Grok 4 Leads at 95.7 Points Despite 38% Instruction Decay Across 15 Models
WDCD Run #360 (2026-10-04) evaluated 15 AI models on multi-turn commitment integrity, with Grok 4 topping the leaderboard at 95.7 points while the coh
Oct 4, 2026
Review GLM-4.6 Surges 26 Points, GPT-5.5 Follows with 10.5 as WDCD Compliance Landscape Shifts
In the WDCD v3.1 pilot, GLM-4.6 rose 26 points to 87.55 and moved into third place, while GPT-5.5 gained 10.5 points; the other 13 evaluated models re
Oct 4, 2026
Review WDCD Five-Scenario Review: Business Rules Trails at 2.13, GPT-6 Series Shows 1.84-Point Imbalance
In WDCD v3.1's five-scenario tests, business rules was the lowest-scoring scenario across all models, with doubao-pro scoring just 2.13/4. GPT-6-sol s
Oct 4, 2026
Review 48.5% Integrity Rate After Three Rounds of Pressure: Grok4 Suffers Zero Collapses, GPT-o3 Collapse Rate Reaches 20.7%
In sampling limited to eight v2 anchor questions, 15 models averaged an R3 integrity rate of just 48.5%, with 41 full collapses out of 435 R3 trials,
Oct 4, 2026
Review Grok 4 Tops WDCD at 95.69 as Qwen3 Max Ranks Last at 73.48, Gap Exceeds 22 Points
Grok 4 scored 95.69 to rank first in the WDCD v3.1 adherence test, while Qwen3 Max ranked last with 73.48, a difference of 22.21 points. The leaderboa
Oct 4, 2026
Lab WDCD Run #346: All 11 Models Hold Zero Instruction Decay, Grok 4 Tops Leaderboard at 100 Points
WDCD Run #346 (2026-09-30) recorded 0% average instruction decay across all 11 tested models, with Grok 4 achieving a perfect 100-point score in the m
Sep 30, 2026
Review Qwen3 Max Surges 20.2 Points, GLM-4.6 Follows Closely, All Five WDCD Models Rise
In this round of WDCD v3.1 testing, Qwen3 Max gained 20.2 points and GLM-4.6 gained 19.9 points, while Claude Sonnet 4.6, Gemini 2.5 Pro, and Grok 4 a
Sep 30, 2026
Review WDCD Five-Scenario Comparative Review: Engineering Standards Lowest at 2.45; Doubao and Claude Bias Gap Reaches 1.45 Points
WDCD v3.1's five-scenario comparative review finds engineering standards to be the weakest area across all models, with Doubao-pro scoring only 2.45/4
Sep 30, 2026
Review 11 Models’ WDCD v2 Anchor R1 Confirmation Rate 0%; Compliance Collapses Completely from the First Round
Across eight v2 anchor questions, all 11 evaluated models had an average R1 confirmation rate of 0/1 (0%). This indicates that hard constraints were n
Sep 30, 2026
Review Grok 4 Scores 100 to Dominate WDCD Commitment-Adherence Leaderboard, Doubao Pro Bottom at 75.3 with 24.7-Point Gap
In the WDCD v3.1 commitment-adherence test, Grok 4 is the only model to score 100.00, while Doubao Pro ranks 11th at 75.30, a 24.7-point gap. The resu
Sep 30, 2026
Review 7-Day Smoke Data: Claude Opus 4.7 Has Highest Trend Score of 25.2, Gemini 3.1 Pro Drops 26.9 Points
Based on Smoke evaluation data from 2026-09-21 to 2026-09-27, Claude Opus 4.7 rose from 70.77 to 96.01 over seven days, with the highest average of 87
Sep 27, 2026
Lab WDCD Run #336: Average Instruction Decay Hits -36.4% as Gemini 3.1 Pro Holds Zero Decay
WDCD Run #336 (2026-09-23) benchmarked 11 models on multi-turn commitment, recording an average instruction decay of -36.4% from Round 1 to Round 3, w
Sep 23, 2026
Review Four Models Slide Collectively in WDCD as Gemini 2.5 Pro Plunges 16.7 Points
In the WDCD v3.1 pilot, Gemini 2.5 Pro fell 16.7 points from Run #331, the largest decline among four evaluated models. DeepSeek V4 Pro, GPT-5.5, and
Sep 23, 2026
Review Data Boundaries Emerge as the Biggest Compliance Blind Spot: 11 Models Score as Low as 1.3, with Gaps up to 2.7
In WDCD v3.1's five constraint scenario tests, data boundary scenarios had the lowest average scores, with Qwen3-Max scoring only 1.3/4 and GLM-4.6 sc
Sep 23, 2026
Review WDCD Three-Round Anchor Test: R3 Integrity Rate 63.6%, GPT-o3 Confirms First Then Collapses
Based on 110 samples from eight v2 anchor questions, 11 models averaged only a 63.6% R3 integrity rate, with three complete collapses. The results sho
Sep 23, 2026
Review Grok 4 Leads WDCD Commitment-Adherence Leaderboard with 93.60; Qwen3 Max Last at 67.30
Grok 4 tops the WDCD commitment-adherence test with 93.60, while Qwen3 Max ranks 11th at 67.30, a 26.3-point gap between top and bottom. The results a
Sep 23, 2026
Lab WDCD Run #331: Grok 4 Leads at 91.8 as Average Instruction Decay Hits -15.1%
WDCD Run #331 (2026-09-20) evaluated 11 models on multi-turn commitment integrity, with Grok 4 topping the ranking at 91.8 points while the cohort ave
Sep 20, 2026
Review GLM-4.6 Plunges 29.8 Points and GPT-5.5 Drops 6 as Both Models Decline Across the Board in WDCD Commitment-Keeping Test
Latest WDCD v3.1 pilot data shows GLM-4.6 dropping 29.8 points and GPT-5.5 dropping 6 points versus Run #326, while none of the other nine evaluated m
Sep 20, 2026
Review WDCD Benchmark Comparison: Data Boundaries Becomes the Hardest Scenario as GLM-4.6 Collapses with Only 1.36
In the WDCD v3.1 test, the data-boundary scenario produced the lowest scores of all five categories, with glm-4.6 scoring just 1.36/4 against 3.8/4 fo
Sep 20, 2026
Review R3 Integrity Rate Only 49.5%: The Compliance Gap Between Grok 4’s Zero Collapse and GLM-4.6’s 27.6% Collapse Rate
In WDCD v2 anchor-task testing, Grok 4 recorded zero complete collapses at R3, while GLM-4.6’s R3 collapse rate reached 27.6%, revealing sharp differe
Sep 20, 2026