Grok 4 Scores 93.80 to Top the Compliance Test, Doubao Pro Trails at 67.30 with a 26.5-Point Gap

In the WDCD v3.1 compliance test, Grok 4 achieved the highest score of 93.80 among the 11 evaluated models, while Doubao Pro trailed at 67.30, a gap of 26.5 points. The v2 anchor questions showed that Grok 4 scored a perfect 1.00/1.00/1.00 across all three rounds R1, R2, and R3, whereas Doubao Pro scored only 0.00/2 in R3.

Ranking Landscape: Top Three and Bottom Gap

The top three—Grok 4 (93.80), GLM-4.6 (92.00), and DeepSeek V4 Pro (90.90)—formed a distinct tier, with a lead of more than 3.8 points over the fourth-place GPT-o3 (87.10). At the bottom, three models—Qwen3 Max (71.00), GPT-5.5 (68.60), and Doubao Pro (67.30)—all scored below 72 points, a gap of over 20 points compared to the top-tier average of 92.23.

The overall full-score rate was 51.8%, and the R3 collapse rate was only 3.6%, indicating that most models maintained constraints under single-round high pressure, but the differences rapidly amplified after multi-round progressive pressure.

Champion Grok 4: The Mechanism Behind Perfect v2 Anchor Scores

Grok 4 scored 1.00/2 in v2 R3 anchor questions, combined with the high retention rate of the S_hold 60-point item in v3 questions, contributing to its total score of 93.80. In both data boundary and safety compliance constraint scenarios, it showed no breach records over 8–12 consecutive conversation rounds. In contrast, the bottom-ranked Doubao Pro scored 0 in the same v2 R3 stage, indicating it lost constraints earliest under authoritative special approval pressure.

Bottom-Ranked Doubao Pro: Recovery Weakness Exposed by Zero R3 Score

Among the 67.30 points of Doubao Pro, the S_recover breach recovery 10-point item and the S_integrity honest self-report 15-point item scored low. The v3 question design required an honest debrief in the final round; Doubao Pro failed to accurately restate the initial constraints after the KBV restatement probe, leading to a significant loss in the S_kbv constraint memory 15-point item. The worst-of-3 sampling further amplified its worst performance.

Top Tier vs. Bottom Gap: Decisive Factor of Pressure Rounds

Although DeepSeek V4 Pro scored a perfect 2.00/2 in R3, its total score still lagged Grok 4 by 2.9 points, mainly due to the S_hold survival score in v3 multi-round progressive pressure. GLM-4.6 increased by 13.7 points compared to the previous evaluation, primarily benefiting from its R2 interference round score recovering from a low level to 1.00.

The bottom three models averaged only 0.67/2 in R3 under the salami slicing and sunk cost progressive pressure scenarios, while the top three averaged 1.33/2. The gap did not stem from single-round collapse but from the attenuation of constraint memory after cumulative multi-round pressure.

Cause Analysis: Interaction Between Constraint Scenarios and Pressure Rounds

Among the five constraint scenarios, resource limitation and engineering specification scenarios had the highest breach rates. The 8–12 round design of v3 questions allowed social identity and authoritative special approval pressures to compound, causing clear differentiation in the S_hold score after the 7th round. The worst-of-3 sampling showed that the top model's worst run still maintained above 85 points, while the bottom model's worst run dropped below 60 points.

Selection Implications: Real Risks of Production Workflow Integration

Enterprises integrating AI into production workflows can prioritize Grok 4 and GLM-4.6 for data boundary and safety compliance scenarios, as their S_hold scores support sustained multi-round constraint retention. In resource limitation scenarios, additional secondary verification guardrails should be set for DeepSeek V4 Pro and GPT-o3, as their R2 interference round scores are 1.00 and 0.00, respectively.

In business rule and engineering specification scenarios, Claude Opus 4.7 (85.80) and Claude Sonnet 4.6 (81.50) can serve as mid-tier options, but human review checkpoints should be added in sunk cost pressure rounds. Doubao Pro and GPT-5.5 scored below 70 points and are not recommended for direct use in production pipelines without human backup.

Strategic Judgment: Signals of Underestimation and Overestimation

GLM-4.6 increased by 13.7 points in this evaluation and achieved a perfect R3 score, suggesting its compliance capability may be underestimated by the market; the next evaluation should focus on whether the S_recover item in v3 questions remains stable. GPT-o3 dropped by 6.9 points this round and scored 0 in R2, indicating decay in its constraint memory under interference rounds; it remains to be seen whether this is a version fluctuation.

Grok 4's current leading margin of 93.80 mainly comes from perfect v2 anchor scores. Whether its S_hold 60-point item can maintain a high level when more parallel hard constraints are added in the next v3 question set will be a key observation point. The low baseline R3 collapse rate of 3.6% implies that the next version of the test may further widen model gaps by increasing the number of rounds.

Compliance capability is not a byproduct of model parameters, but a fundamental threshold that determines production usability.

Data Source: YZ Index WDCD Compliance Leaderboard | Run #242 · Overall Ranking | Evaluation Methodology

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!