Grok 4 Tops WDCD with 94.80 Points, Qwen3 Max Ranks Last with 69.20, a 25.6-Point Gap

In the WDCD v3.1 compliance test, Grok 4 ranked first with 94.80 points, while Qwen3 Max placed 11th with 69.20 points, a total score gap of 25.6 points. The result is derived from the equally weighted average of 10 v3 questions and 8 v2 anchor questions, using worst-of-3 sampling.

Ranking Landscape and Core Data

The top four — Grok 4 (94.80), Gemini 3.1 Pro (91.30), GLM-4.6 (88.50), and Claude Opus 4.7 (85.40) — form a clear leading tier. Fifth-place GPT-o3 (83.60) trails sixth-place DeepSeek V4 Pro (83.10) by just 0.5 points, yet their R3 scores differ significantly: GPT-o3 scored 0.00/2 on R3 while DeepSeek V4 Pro scored 1.50/2. At the tail end, Qwen3 Max's 69.20 points sits 5.6 points below tenth-place Gemini 2.5 Pro (74.80).

The overall full-score rate is 55.5%, with an R3 collapse rate of 4.5%. This suggests that most models can uphold their commitments during the first two rounds of interference, but nearly half show varying degrees of constraint breakdown when entering the high-pressure third round.

R3 Pressure Round Determines Final Rankings

Differences in overall compliance scores mainly stem from the continuous pressure phase of v3 questions and the R3 segment of v2 anchor questions. Both Grok 4 and Gemini 3.1 Pro scored 1.50/2 on R3, demonstrating they retain at least 75% of their initial constraints under the four-tier escalation of social proof, authority-granted exceptions, salami slicing, and sunk cost. GPT-o3's 0.00/2 on R3 means it completely lost its constraints within the same pressure sequence, resulting in heavy deductions on the S_hold compliance survival metric.

GLM-4.6 scored only 1.00/2 on R3 yet still ranks third, relying mainly on stable outputs of 15 points on the S_kbv constraint memory metric and 10 points on S_recover post-breakdown recovery for v3 questions. Qwen3 Max stayed at 1.00/2 on R3 throughout and triggered multiple 0-point penalties on the S_integrity honest self-report metric under worst-of-3 sampling, dragging its total down to 69.20.

Performance Differences Across Five Constraint Scenarios

Data security boundary and security compliance scenarios have the greatest impact on R3 scores. Grok 4 maintained constraints through rounds 8-10 in worst-of-3 testing across both scenario types, with S_hold scores approaching the 60-point ceiling. Qwen3 Max showed constraint breakdown as early as round 5 in resource limitation scenarios, leaving S_hold at around 30 points.

Engineering specification scenarios are relatively more forgiving, with the 11 models averaging 1.20/2 on R3. However, in business rule scenarios, both GPT-o3 and GPT-5.5 dropped to 0.50/2 as early as R2, indicating that early interference can already shake constraint anchors.

Practical Implications for Production Pipeline Integration

Enterprises integrating AI into production pipelines can reference the following assessments: Grok 4 and Gemini 3.1 Pro, with R3 scores of 1.50/2 in data boundary and security compliance scenarios, can be deployed directly under low-guardrail conditions; GPT-o3, given its 0.00/2 R3 score, requires an additional external constraint validation layer in any customer service or content generation workflow involving multi-round user pressure.

DeepSeek V4 Pro fell 8.0 points this cycle compared with the previous one, primarily due to R2 interference performance dropping from 1.00 to 0.50, indicating weakened resistance to continuous contextual interference. Enterprises that have already deployed this model based on prior-cycle data should reassess guardrail strength for resource-limited tasks in the current cycle.

Strategic Judgments and Next-Cycle Verification Signals

Claude Opus 4.7 ranks fourth with 85.40 points, and its R3 score of 1.50/2 outperforms the same family's Claude Sonnet 4.6 (79.00 points, R3 1.00/2), revealing a generational gap in compliance capability under high-pressure scenarios between models of different sizes from the same vendor. The market may be overestimating the actual constraint retention capability of the Sonnet series.

GPT-o3 dropped 11.6 points this cycle with R3 falling directly to zero — a signal worth close attention in the next cycle: whether the newly added KBV recitation probes in v3 questions triggered earlier breakdowns. Qwen3 Max's last-place finish likely stems from multiple zero scores on the S_integrity metric, suggesting it tends to falsely claim innocence rather than honestly acknowledge failures after breaking constraints.

Doubao Pro rose 13.8 points this cycle, and GLM-4.6 rose 9.9 points, both benefiting from R3 improvements from 1.00 to higher levels. Should this trend continue in the next test cycle, both models could enter the top five.

Compliance capability is not a byproduct of model parameters; it is a genuine moat for production deployment. The gap between Grok 4's 94.80 points and Qwen3 Max's 69.20 points has already drawn that moat on enterprise model selection checklists.

Data source: YZ Index WDCD Compliance Leaderboard | Run #276 · Overall Ranking | Evaluation Methodology

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!