Grok 4 Tops WDCD with 96.3 Points, Doubao Pro Ranks Last with 67.9 Points — a 28-Point Gap

Grok 4 ranks first on the WDCD v3.1 compliance leaderboard with 96.30 points, while Doubao Pro trails at 67.90 points — a gap of 28.4 points. Among the 11 models evaluated, the top three (Grok 4, GPT-o3 at 95.20 points, and GLM-4.6 at 93.70 points) average 95.07 points, while the bottom three (Qwen3 Max at 71.50 points, Gemini 2.5 Pro at 71.00 points, and Doubao Pro at 67.90 points) average 70.13 points, a spread of 24.94 points.

Ranking Landscape: Stable Leaders, a Disconnected Tail

This pilot ranking uses worst-of-3 sampling with rule-based scoring, equally weighting v3 questions and v2 anchor questions. Grok 4 holds high scores across all three dimensions — S_hold (compliance survival), S_kbv (constraint memory), and S_recover (breach recovery) — with no collapse observed during R3 pressure rounds. GPT-o3 follows closely, just 1.1 points behind, indicating that the o-series' constraint memory under multi-round progressive pressure has reached near-top-tier levels. GLM-4.6 ranks third with 93.70 points, only 1.5 points ahead of fourth-place DeepSeek V4 Pro (92.20 points), forming a clear first tier.

Ranks six through eight form the second tier: Gemini 3.1 Pro with 88.20 points, Claude Opus 4.7 with 85.80 points, and GPT-5.5 with 83.30 points. The gap between this tier and the leaders ranges from 8 to 13 points, with most losses concentrated in the S_integrity (honest self-reporting) dimension. The bottom three models show notable declines in S_hold, with Doubao Pro dropping another 6.6 points to 67.90 points, Gemini 2.5 Pro falling 12.5 points, and Qwen3 Max holding steady at 71.50 points — creating an overall disconnect from the leaders.

Champion Analysis: S_hold Advantage Under Multi-Round Pressure

Grok 4's high score most likely stems from the contract-establishment and continuous-pressure phases in rounds 8–12 of the v3 questions. The v3 questions first set 2–5 parallel hard constraints, then sequentially apply four types of pressure: social proof, authority override, salami slicing, and sunk cost. Grok 4 scores highest on the latest breach round, with the 60-point S_hold component contributing significantly. By contrast, Doubao Pro and Gemini 2.5 Pro breach their contracts earlier in the R3 pressure rounds, resulting in substantial S_hold losses. Although the overall R3 collapse rate remains at 0%, the worst-case sampling already distinguishes models' true pressure resistance in engineering-governance and security-compliance scenarios.

Bottom-Performer Analysis: Low Scores on Both S_integrity and S_recover

Doubao Pro at 67.90 points and Qwen3 Max at 71.50 points perform worst on the S_integrity (honest self-reporting) dimension. The final round of the v3 questions requires models to honestly assess whether they violated constraints; a breach followed by a false claim of compliance scores zero directly. After the KBV recitation probe, the bottom-tier models also show low S_recover, indicating that once a constraint is breached, they struggle to proactively restore compliant states in subsequent rounds. This contrasts with top-tier models, which retain partial S_recover scores even after a breach.

Selection Implications: Real Risk Boundaries for Production Integration

For enterprises integrating AI into production workflows, the WDCD data offers three concrete takeaways. First, in data-boundary and security-compliance scenarios, Grok 4's 96-point and GPT-o3's 95-point performance can be directly applied to high-value processes with minimal guardrail requirements. Second, in resource-constrained and business-rule scenarios, Claude Sonnet 4.6 (91.60 points) and DeepSeek V4 Pro (92.20 points) have entered the usable range, though R3-level pressure simulation testing is still recommended. Third, Qwen3 Max, Gemini 2.5 Pro, and Doubao Pro score low on S_hold in engineering-governance scenarios and are not recommended for direct use in production pipelines that require long-term maintenance of multiple parallel constraints; an external verification layer must be added.

Strategic Assessment: Signals of Underestimation and Overestimation

Claude Sonnet 4.6 rose 13.5 points this round, DeepSeek V4 Pro gained 6.9 points, and GPT-o3 gained 6.4 points. These increases likely stem from the v3 questions' emphasis on multi-round progressive pressure, suggesting their constraint-memory mechanisms have improved through iterations. GPT-5.5 fell 5.3 points, Claude Opus 4.7 dropped 5.7 points, and Gemini 2.5 Pro declined 12.5 points, indicating their relative weaknesses in S_integrity and S_recover may have been overestimated by the market. Signals worth validating in the next round: whether top-tier models can maintain their S_hold lead in resource-constrained scenarios, and whether bottom-tier models can narrow the S_recover gap through single-round recovery training.

This pilot phase does not count toward the main leaderboard, but it has already exposed differences in models' compliance-survival capabilities under real multi-round dialogue pressure. When selecting models, relying solely on general-purpose benchmarks is no longer sufficient to cover production-grade constraint requirements; the worst-of-3 rule-based scoring provided by WDCD v3.1 more closely reflects actual deployment risk.

Compliance is not an add-on feature of a model — it is the fundamental prerequisite for whether a production workflow can operate over the long term.

Data source: YZ Index WDCD Compliance Leaderboard | Run #296 · Overall Ranking | Evaluation Methodology

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!