Grok 4 Tops WDCD Compliance Leaderboard with 94.80 Points, Doubao Pro Trails at 64.20 Points, a 30-Point Gap

In the WDCD v3.1 compliance test, Grok 4 ranked first with 94.80 points, while Doubao Pro placed 11th with 64.20 points, a difference of 30.6 points.

Ranking Landscape and Core Data

This test covered 11 models, with a question pool consisting of 17 v3 multi-round progressive pressure questions and 8 v2 anchor questions. The overall compliance score is calculated as the equal-weighted average of the v3 native percentage score and the v2 anchor question score after conversion, using a worst-of-3 sampling method. Grok 4 achieved R1=1.00, R2=1.00, R3=1.50/2, indicating it only lost 0.5 points in the final pressure round of the three anchor questions. DeepSeek V4 Pro followed closely with 93.60 points, also scoring R3=1.50/2. GLM-4.6 and Claude Opus 4.7 scored 93.50 and 92.60 points respectively, with R3 scores of 1.00/2 or 1.50/2.

The gap is significant among tail-end models. GPT-5.5 scored R3=0.00/2, Gemini 2.5 Pro and Qwen3 Max also had R3 scores of 0.00/2 or 1.50/2 but with overall lower scores. Doubao Pro had an R1 of only 0.50, indicating notable degradation as early as the first round of constraint injection.

Root Cause Analysis: R3 Pressure Round Determines the Gap

Data shows that the R3 pressure round is the primary mechanism for score differentiation. The top four models averaged 1.375/2 on R3, while the bottom four averaged only 0.625/2. In the design of v3 questions, R3 corresponds to the final level of continuous pressure, including sunk costs and authoritative overrides stacked together. In data boundary and safety compliance scenarios, models must maintain multiple parallel hard constraints simultaneously; low R3 scores directly drag down the S_hold compliance survival score.

Claude Opus 4.7 improved by 6.8 points this round, mainly due to R3 recovering from a low in the previous round to 1.50/2, indicating enhanced resilience under salami-slicing progressive pressure. Gemini 3.1 Pro dropped by 5.6 points, corresponding to R3 falling from a possibly higher level to 1.00/2, suggesting faster decay of constraint memory after multi-round KBV paraphrase probes.

Selection Implications for Production Pipeline Integration

Enterprises integrating AI into production pipelines can delineate usage boundaries based on WDCD scores. Grok 4 and DeepSeek V4 Pro, with higher R3 scores in engineering specification and resource-constrained scenarios, are suitable for direct integration into processes requiring long-term maintenance of business rules, such as automated contract clause verification or API quota control. Claude Opus 4.7, with R3=1.50/2, can be a first choice in safety compliance scenarios but still requires additional secondary verification in data boundary scenarios.

Tail-end models like Doubao Pro and Qwen3 Max, with R3=0.00/2 or degraded R1, show significantly higher probability of compliance failure in business rule and safety compliance scenarios. If enterprises must use these models, they should add hard guardrails at the integration layer, such as enforcing constraint paraphrasing before output or setting up independent rule engines to intercept.

Strategic Assessment

Current data may underestimate the recovery potential of the Claude series under multi-round pressure. Both Sonnet 4.6 and Opus 4.7 improved by over 6 points this round, warranting continued observation of R3 stability in the next round. GPT-5.5 scored R3=0.00/2 while GPT-o3 scored R3=0.50/2, showing significant variance between different versions from the same vendor, highlighting the non-linear impact of version iteration on compliance capability.

Top models perform more stably during the KBV paraphrase probe phase of v3 questions, contributing higher S_kbv and S_recover scores. Tail-end models tend to lose points earlier in the S_hold phase. If the next round of v3 questions increases the number of parallel constraints in engineering specification scenarios, the advantage of models currently with high R3 scores may further widen.

Compliance capability is not a byproduct of model parameters but a core threshold determining whether a model can remain in the production pipeline long-term.

Data source: YZ Index WDCD Compliance Leaderboard | Run #253 · Overall Ranking | Evaluation Methodology

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!