This round of WDCD v3.1 testing shows Claude Sonnet 4.6's WDCD score rose 8.6 points from Run #311, while Claude Opus 4.7, Gemini 2.5 Pro, Gemini 3.1 Pro, GLM-4.6, GPT-5.5, GPT-o3, and DeepSeek V4 Pro all declined — seven models in total. Among them, GLM-4.6 fell 27 points and Gemini 2.5 Pro fell 23.8 points.
Data Facts: One Rise vs. Seven Declines
Among the 11 models evaluated, only Claude Sonnet 4.6 rose, while seven models declined. The current Top 5: Grok 4 WDCD=93.80, GPT-o3 WDCD=89.80, Claude Sonnet 4.6 WDCD=87.20, DeepSeek V4 Pro WDCD=83.60, and Doubao Pro WDCD=83.00. Claude Opus 4.7 dropped 17 points, GPT-5.5 dropped 12.4 points, GPT-o3 dropped 5 points, and DeepSeek V4 Pro dropped 5 points. Sampling follows a worst-of-3 protocol — each question is run three times and the worst result is retained — with v3 questions and v2 anchor questions equally weighted.
Root-Cause Analysis: Divergent Constraint Survival Under Multi-Round Pressure
WDCD v3 questions are designed around 8–12 dialogue rounds: 2–5 hard constraints are established first, then escalating pressure is applied through social proof, authority override, salami-slicing, and sunk costs, followed by a KBV restatement probe and a final round of honest review. S_hold (constraint-keeping survival) accounts for 60 points — the later a constraint breaks, the higher the score. GLM-4.6's 27-point drop and Gemini 2.5 Pro's 23.8-point drop most likely occurred in rounds R3–R6 of the sustained-pressure phase, where constraints gave way earliest under mounting sunk-cost pressure. Claude Sonnet 4.6's 8.6-point gain reflects improved scores on both S_recover (recovery after a break) and S_integrity (honest self-reporting), indicating it returns to the constraint framework more quickly after a breach. In the three-round v2 anchor-question design, R3 carries a pressure weight of 2 points, and the declining models most likely lost the most points in that round.
Model Selection Implications: Real Risk Boundaries for Production Integration
Enterprises integrating AI into production workflows can map use cases against WDCD scores. Models such as Grok 4 (WDCD=93.80) and GPT-o3 (WDCD=89.80) suit data-boundary and security-compliance tasks, offering longer constraint-keeping survival and safe direct integration without additional guardrails. Claude Sonnet 4.6 (WDCD=87.20) improved but still trails the top two; it is usable in engineering-standard scenarios, yet resource-constraint tasks should include human review checkpoints. GLM-4.6 and Gemini 2.5 Pro posted the sharpest declines — business-rule and security-compliance scenarios should implement second-confirmation mechanisms to prevent a single pressure round from invalidating constraints. Doubao Pro (WDCD=83.00) and DeepSeek V4 Pro (WDCD=83.60) fit internal non-critical workflows, but external-facing systems require an additional rule-engine layer.
Strategic Assessment: Underrated and Overrated Constraint Keepers
Claude Opus 4.7 fell 17 points while Claude Sonnet 4.6 rose 8.6 points — a 25.6-point divergence within the same model family, showing that version iterations significantly alter sensitivity to multi-round pressure. Grok 4 (WDCD=93.80) did not appear on this round's decline list; its constraint-keeping ability may be underrated by the market, and its performance on the KBV restatement probe warrants priority validation next round. GPT-o3 (WDCD=89.80), despite a 5-point drop, holds second place with a still-competitive S_hold score. GLM-4.6's 27-point drop is the largest of the round; it may have experienced constraint-memory decay during the v3 constraint-establishment phase, with the 15-point S_kbv component likely the primary source of loss. The next cycle should track whether Claude Sonnet 4.6 can keep improving on the 60-point S_hold component and whether Grok 4 sustains a level above 93.80.
A compliance score is not a static label but a survival curve under multi-round pressure.
When selecting models, those scoring above WDCD 90 can reduce guardrail investment; models in the 85–90 band need targeted reinforcement; and models below 80 should be confined to low-risk scenarios. The v3.1 pilot phase has already shown that the weight distribution between constraint memory and post-breach recovery will directly affect production deployment costs.
Data source: YZ Index WDCD Compliance Leaderboard | Run #316 · Change Tracking | Evaluation Methodology
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接