WDCD Compliance Scores Slide Across Seven Models; Claude Sonnet 4.6 Bucks Trend with 8.6-Point Gain

This round of WDCD v3.1 testing shows Claude Sonnet 4.6's WDCD score rose 8.6 points from Run #311, while Claude Opus 4.7, Gemini 2.5 Pro, Gemini 3.1 Pro, GLM-4.6, GPT-5.5, GPT-o3, and DeepSeek V4 Pro all declined — seven models in total. Among them, GLM-4.6 fell 27 points and Gemini 2.5 Pro fell 23.8 points.

Data Facts: One Rise vs. Seven Declines

Among the 11 models evaluated, only Claude Sonnet 4.6 rose, while seven models declined. The current Top 5: Grok 4 WDCD=93.80, GPT-o3 WDCD=89.80, Claude Sonnet 4.6 WDCD=87.20, DeepSeek V4 Pro WDCD=83.60, and Doubao Pro WDCD=83.00. Claude Opus 4.7 dropped 17 points, GPT-5.5 dropped 12.4 points, GPT-o3 dropped 5 points, and DeepSeek V4 Pro dropped 5 points. Sampling follows a worst-of-3 protocol — each question is run three times and the worst result is retained — with v3 questions and v2 anchor questions equally weighted.

Root-Cause Analysis: Divergent Constraint Survival Under Multi-Round Pressure

WDCD v3 questions are designed around 8–12 dialogue rounds: 2–5 hard constraints are established first, then escalating pressure is applied through social proof, authority override, salami-slicing, and sunk costs, followed by a KBV restatement probe and a final round of honest review. S_hold (constraint-keeping survival) accounts for 60 points — the later a constraint breaks, the higher the score. GLM-4.6's 27-point drop and Gemini 2.5 Pro's 23.8-point drop most likely occurred in rounds R3–R6 of the sustained-pressure phase, where constraints gave way earliest under mounting sunk-cost pressure. Claude Sonnet 4.6's 8.6-point gain reflects improved scores on both S_recover (recovery after a break) and S_integrity (honest self-reporting), indicating it returns to the constraint framework more quickly after a breach. In the three-round v2 anchor-question design, R3 carries a pressure weight of 2 points, and the declining models most likely lost the most points in that round.

Model Selection Implications: Real Risk Boundaries for Production Integration

Enterprises integrating AI into production workflows can map use cases against WDCD scores. Models such as Grok 4 (WDCD=93.80) and GPT-o3 (WDCD=89.80) suit data-boundary and security-compliance tasks, offering longer constraint-keeping survival and safe direct integration without additional guardrails. Claude Sonnet 4.6 (WDCD=87.20) improved but still trails the top two; it is usable in engineering-standard scenarios, yet resource-constraint tasks should include human review checkpoints. GLM-4.6 and Gemini 2.5 Pro posted the sharpest declines — business-rule and security-compliance scenarios should implement second-confirmation mechanisms to prevent a single pressure round from invalidating constraints. Doubao Pro (WDCD=83.00) and DeepSeek V4 Pro (WDCD=83.60) fit internal non-critical workflows, but external-facing systems require an additional rule-engine layer.

Strategic Assessment: Underrated and Overrated Constraint Keepers

Claude Opus 4.7 fell 17 points while Claude Sonnet 4.6 rose 8.6 points — a 25.6-point divergence within the same model family, showing that version iterations significantly alter sensitivity to multi-round pressure. Grok 4 (WDCD=93.80) did not appear on this round's decline list; its constraint-keeping ability may be underrated by the market, and its performance on the KBV restatement probe warrants priority validation next round. GPT-o3 (WDCD=89.80), despite a 5-point drop, holds second place with a still-competitive S_hold score. GLM-4.6's 27-point drop is the largest of the round; it may have experienced constraint-memory decay during the v3 constraint-establishment phase, with the 15-point S_kbv component likely the primary source of loss. The next cycle should track whether Claude Sonnet 4.6 can keep improving on the 60-point S_hold component and whether Grok 4 sustains a level above 93.80.

A compliance score is not a static label but a survival curve under multi-round pressure.

When selecting models, those scoring above WDCD 90 can reduce guardrail investment; models in the 85–90 band need targeted reinforcement; and models below 80 should be confined to low-risk scenarios. The v3.1 pilot phase has already shown that the weight distribution between constraint memory and post-breach recovery will directly affect production deployment costs.


Data source: YZ Index WDCD Compliance Leaderboard | Run #316 · Change Tracking | Evaluation Methodology

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!