WDCD Compliance Leaderboard: Grok 4 Wins with 91.04 Points, Doubao Pro Trails at 58—a 33-Point Gap

In the WDCD v3.1 compliance test, Grok 4 took first place with 91.04 points, while Doubao Pro ranked last with 58.00 points—a 33.04-point gap between the top and bottom. The results were derived from 11 models' worst-of-3 sampling across 25 questions, with v3 questions and v2 anchor questions equally weighted.

Ranking Landscape: R3 Pressure Round Determines the Gap

The top two models, Grok 4 (91.04) and DeepSeek V4 Pro (89.04), scored 1.25/2 and 1.50/2 respectively on R3, significantly ahead of the mid-tier. Claude Sonnet 4.6 (83.72), GPT-o3 (83.40), and Gemini 3.1 Pro (83.36) posted R3 scores of 1.00/2, 1.00/2, and 0.88/2 respectively, with the gap mainly stemming from R3 performance after R2 interference. GPT-5.5 managed only 0.50/2 on R3, landing at 75.04 overall; Doubao Pro scored 0.00/2 across R1, R2, and R3, dragging it directly down to 58.00.

The overall perfect score rate was 49.5%, with an R3 collapse rate of 12.4%, indicating that more than one in ten models failed to honor their commitments after four levels of escalating pressure: continuous social proof, authority approval, salami slicing, and sunk cost. In v2 anchor questions, R3 carries the highest weight (2 points), so R3 scores directly amplify the overall differences.

Champion's Edge: Grok 4's R3 Performance Under Multi-Round Escalating Pressure

Grok 4 scored R1=1.00, R2=0.88, and R3=1.25/2, with a balanced mix of the four sub-components—S_hold, S_kbv, S_recover, and S_integrity—on v3 questions. Maintaining a high survival rate under parallel hard constraints in the R3 phase shows stable performance across KBV paraphrase probes and final-round honesty reviews. DeepSeek V4 Pro reached 1.50/2 on R3 but only 0.63 on R2, indicating it is slightly weaker than Grok 4 during the interference phase.

Claude Sonnet 4.6 rose 8.8 points this cycle, primarily driven by R2 recovering from a low base to 0.88, with R3 holding at 1.00/2. GPT-5.5 gained 5.0 points, but R3 remains stuck at 0.50/2, suggesting limited recovery capability. Doubao Pro scored zero across all three rounds, indicating it cannot establish effective constraint memory even at the commitment stage.

Tail-End Gap: Constraint Scenario Failures for Doubao Pro and Qwen3 Max

Doubao Pro experienced R1-level breakdowns across all five scenario types—data boundaries, resource limits, business rules, security compliance, and engineering standards—with S_hold scores near zero. Qwen3 Max managed only 0.63/2 on R3, 0.37–0.62 points below the mid-tier. Gemini 2.5 Pro scored 1.00/2 on R3 but finished at 71.96 overall, showing that the cumulative effect of multi-round pressure in v3 questions is pronounced.

These gaps are not random. The v3 questions are designed with 8–12 rounds of dialogue, with pressure escalating level by level. The 12.4% R3 collapse rate directly corresponds to where tail-end models break their commitments in real conversations.

Implications for Production Integration Choices

When enterprises integrate AI into high-constraint scenarios, the 91- and 89-point levels of Grok 4 and DeepSeek V4 Pro can reduce the need for additional guardrails. In data boundary and security compliance scenarios, models scoring above 1.25/2 on R3 can still maintain constraints after sunk-cost pressure and are suitable for direct production use. Models with R3 scores below 0.75/2 (such as GPT-5.5, Qwen3 Max, and Doubao Pro) require an independent validation layer in business rule and engineering standard scenarios; otherwise, they risk violating hard constraints under sustained pressure.

Claude Sonnet 4.6 and GPT-o3 fall in the 83-point range, suitable for low-to-mid-risk processes, but resource-limited scenarios still require monitoring of post-R2 recovery performance.

Strategic Assessment: Signals of Underrated and Overrated Compliance Capabilities

The mere 2-point gap between Grok 4's 91.04 and DeepSeek V4 Pro's 89.04 shows that top-tier models have already formed a distinct echelon under v3's multi-round pressure. Claude Sonnet 4.6's single-cycle gain of 8.8 points suggests its R2 recovery mechanism may continue to improve in the next version, making it worth focused validation next cycle. GPT-5.5's R3 score of only 0.50/2 despite a 5-point overall gain indicates room for improvement in its S_recover and S_integrity sub-components, but the R3 pressure round remains the bottleneck.

Doubao Pro's −16.0-point decline and zero scores across all rounds suggest its compliance capability is overvalued by the market under the current v3.1 question pool. Qwen3 Max's 67.68 points—more than 15 points behind the mid-tier—likewise requires additional guardrails before entering production. The next cycle should focus on observing the S_hold survival curves of models scoring 1.00/2 or above on R3 in security compliance scenarios.

Compliance is not a marketing slogan—it is the real score after the R3 pressure round.

Data source: YZ Index WDCD Compliance Leaderboard | Run #271 · Overall Ranking | Evaluation Methodology

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!