Instruction Compliance & WDCD

170 articles · Page 1 of 9
Does your AI model actually follow instructions? Instruction compliance is the most critical evaluation dimension for enterprise AI deployment, yet traditional benchmarks rarely test it. WDCD (Winzheng Dynamic Contextual Decay) is the world's first systematic test measuring how AI models' commitment to instructions decays over extended dialogue — using three rounds of 2,000-5,000 word professional distractions across 30 constraint questions in 5 real-world scenarios, with 100% rule-based scoring and zero AI judges. The YZ Index Integrity Rating also deploys 42 canary probes to detect fabricated citations and hallucinated data. This topic covers instruction compliance research, hallucination detection methods, and WDCD test result analysis.

In-depth Guides

Review We Crafted Four Meaningless Rules to Trick AI into Violating Them — and Failed on Every Count
In a deliberately designed sting operation, models across three tiers were handed four meaningless rules and put through seven rounds of social-engine
Aug 21, 2026
Review Gemini 2.5 Pro Integrity Rating Fails: Probe 25, Main Leaderboard 18.37, Code Execution 23.50
Gemini 2.5 Pro received a failing integrity rating in the 2026-08-21 Run#288 smoke test, with a probe score of only 25.00, a main leaderboard score of
Aug 21, 2026
Lab WDCD Run #285: Average Instruction Decay Hits 54.5% Across 11 Models, Grok 4 Leads with Zero Drift
WDCD Run #285 (2026-08-19) tested 11 frontier models across three dialogue rounds and recorded an average commitment decay of 54.5%, with Grok 4 toppi
Aug 19, 2026
Review WDCD v3.1 Cycle: Gemini 2.5 Pro Up 8.7 Points, Doubao Pro the Only Decliner at 7.3
WDCD v3.1 pilot data shows Gemini 2.5 Pro up 8.7 points this cycle as one of four risers, while Doubao Pro fell 7.3 points as the sole decliner. Grok
Aug 19, 2026
Review WDCD Comparative Review: Security Compliance Lowest at 1.15, 2.85-Point Compliance Gap Across 11 Models
In the WDCD v3.1 five-scenario comparative review, security compliance proved to be the weakest scenario for instruction adherence across all models,
Aug 19, 2026
Review After Three Rounds of Pressure, R3 Integrity Rate Falls to Just 22.7%: A Record of 11 Models' WDCD Compliance Collapse
Across three rounds of testing on 8 v2 anchor questions, 11 models posted a 100% R1 confirmation rate and 91% R2 resistance, but R3 integrity plunged
Aug 19, 2026
Review Grok 4 Tops WDCD Promise-Keeping Leaderboard with 97.5 Points, Qwen3 Max Ranks Last with 69.5 Points, a 28-Point Gap
In the WDCD v3.1 promise-keeping test, Grok 4 ranked first with 97.50 points, while Qwen3 Max ranked 11th with 69.50 points, a 28-point gap between th
Aug 19, 2026
Lab WDCD Run #276: Grok 4 Leads with 94.8 Points as Average Instruction Decay Hits -18.2%
WDCD Run #276 (2026-08-12) evaluated 11 models on multi-turn commitment integrity, with Grok 4 taking the top spot at 94.8 points while the fleet-wide
Aug 12, 2026
Review WDCD v3.1 Cycle Tracking: Doubao Pro +13.8, GPT-o3 -11.6, Major Reshuffle in Commitment Rankings
In WDCD v3.1 pilot testing, Doubao Pro climbed 13.8 points over Run #271 while GPT-o3 fell 11.6 points, triggering a major reshuffle in the commitment
Aug 12, 2026
Review Business Rules Scenario Lowest at Just 1.55 Points; Claude Crashes to 2 Points in Resource Constraints
In the WDCD v3.1 five-scenario evaluation, the Business Rules track proved the hardest, with qwen3-max scoring only 1.55/4, while claude-opus-4.7 fell
Aug 12, 2026
Review R3 Integrity Rate Only 59.1%: GPT-o3's 20% Collapse Rate Exposes Three-Round Compliance Gap
In the WDCD v3.1 pilot, the average performance of 11 models on 8 v2 anchor questions exhibited a clear three-round degradation: R1 confirmation rate
Aug 12, 2026
Review Grok 4 Tops WDCD with 94.80 Points, Qwen3 Max Ranks Last with 69.20, a 25.6-Point Gap
In the WDCD v3.1 compliance test, Grok 4 ranked first with 94.80 points while Qwen3 Max placed last with 69.20 points, a 25.6-point gap. The R3 pressu
Aug 12, 2026
Review GLM-4.6 Code Execution Drops 12.5 Points, Perfect Material Constraint Score Lifts Main Leaderboard by 5.4
GLM-4.6 scored 79.38 on today's Smoke evaluation main leaderboard, with code execution falling 12.5 points to 62.50 while material constraints reached
Aug 12, 2026
Review GLM-4.6 Integrity Rating Drops from Pass to Fail, Task Expression Plunges 25 Points While Main Leaderboard Rises 12.8
In today's Smoke evaluation, GLM-4.6's integrity rating dropped from pass to fail, and its Task Expression score fell sharply from 45.00 to 20.00. The
Aug 11, 2026
Lab WDCD Run #271: Grok 4 Leads as Average Instruction Decay Drops to 0.9%
WDCD Run #271 (2026-08-09) tested 11 models across three rounds of multi-turn commitment scoring, recording an average instruction decay of just 0.9%
Aug 9, 2026
Review Claude Sonnet 4.6 Up 8.8 Points, Doubao Pro Down 16 Points: WDCD v3.1 Shows Commitment-Keeping Divergence
Claude Sonnet 4.6 rose 8.8 points to 83.72 on the WDCD v3.1 benchmark while Doubao Pro fell 16 points, as the evaluation begins to differentiate model
Aug 9, 2026
Review WDCD Cross-Review: Data Boundaries Lowest Across All Scenarios, 11 Models Average Only 2.8, doubao-pro Collapses to 1.4
In the WDCD v3.1 five-scenario benchmark, data boundaries scored the lowest of all scenarios, with 11 models averaging just around 2.8. doubao-pro was
Aug 9, 2026
Review WDCD Three-Round Anchors: Doubao Pro Collapses 32% of the Time While Grok Has Zero Collapses; 34 Zero Scores Expose Cracks in Constraint Adherence
In a sample limited to eight v2 anchor questions, 11 models recorded 34 complete R3 collapses out of 275 tests. Doubao Pro showed a 32% R3 collapse ra
Aug 9, 2026
Review WDCD Compliance Leaderboard: Grok 4 Wins with 91.04 Points, Doubao Pro Trails at 58—a 33-Point Gap
In the WDCD v3.1 compliance test, Grok 4 ranked first with 91.04 points while Doubao Pro placed last with 58.00 points, a 33.04-point gap between the
Aug 9, 2026
Review Claude Sonnet 4.6 Drops 22.6 Points, GLM-4.6 Fluctuates 59.9, GPT-o3 Rises 6.7 Points — Smoke Weekly Trend
During August 6–9, 2026, Claude Sonnet 4.6 posted the largest single-model decline in the Smoke rapid test, falling 22.6 points, while GPT-o3 rose 6.7
Aug 9, 2026