97.7 vs 75.2: DeepSeek V4 Pro Dominates Qwen3 Max in Adherence with 22-Point Gap

DeepSeek V4 Pro ranks first on the WDCD v3.1 Adherence Overall Leaderboard with 97.70 points, while Qwen3 Max ranks 11th with 75.20 points, leaving a 22.5-point gap between the top and the tail.

Ranking Landscape: Top Three Clustered, Bottom Three Fall Away

Among the 11 models in this round, DeepSeek V4 Pro, Gemini 3.1 Pro, and Grok 4 form the first tier, scoring 97.70, 96.70, and 96.00, respectively. All three maintain full marks in R1 and R2, with only a 0-1 point difference in R3. Fourth through seventh places—Gemini 2.5 Pro (94.50), GPT-o3 (94.20), Claude Opus 4.7 (92.60), and GPT-5.5 (90.10)—make up the second tier, with 0-point records already appearing in R2. After eighth place, Claude Sonnet 4.6 (83.60), Doubao Pro (76.70), GLM-4.6 (76.50), and Qwen3 Max (75.20) open a clear gap.

Why the Champion Won: Complete Survival Under Multi-Turn Pressure in v3

DeepSeek V4 Pro scored highest on the S_hold adherence survival dimension for v3 tasks and was the last to be broken. The v3 task design includes 8-12 rounds of dialogue: it first establishes 2-5 hard constraints, then applies social proof, special authority approval, salami slicing, and sunk-cost pressure in sequence, and finally runs a KBV recall probe and an honest final-round review. This model completed the entire R1-R3 process without being broken in all 5 constraint scenarios (data boundaries, resource limits, business rules, security compliance, and engineering standards), and earned full marks on S_kbv constraint memory (15 points), S_recover breach recovery (10 points), and S_integrity honest self-reporting (15 points).

Bottom-Tier Mechanism: Chain Reaction from Failure at R1

GLM-4.6 and Qwen3 Max scored 0 in the R1 stage of the v2 anchor task, meaning they were breached as early as the initial constraint injection stage. GLM-4.6 scored 0 across all three anchor-task rounds, R1, R2, and R3, and its S_hold score on v3 tasks fell sharply as a result. Although Qwen3 Max recovered 2 points in R3, its earlier 0 in R2 had already dragged down its total score. The sampling standard is worst-of-3: each task is run 3 times and the worst result is taken. These two models failed at the R1 stage in at least one test round.

Selection Implications of the Top-to-Bottom Gap

For enterprises integrating AI into production workflows, the gap between WDCD scores of 97.70 and 75.20 directly corresponds to different risk exposures. DeepSeek V4 Pro can still maintain constraints in the R3 stage under security compliance and engineering standards scenarios, making it suitable for direct integration into internal approval workflows that require long-term state retention. Qwen3 Max and GLM-4.6 fail at the initial constraint injection stage, meaning that in the same scenarios they must be paired with additional external guardrails or human review; otherwise, a single conversation can breach data boundaries or resource limits.

Practical Boundaries of Second-Tier Models

Gemini 2.5 Pro and GPT-o3 already scored 0 in R2, indicating that their adherence weakens when facing sustained social proof or special authority approval pressure. If enterprises plan to use these two models for business-rule tasks, they need to add secondary confirmation checkpoints in the dialogue rounds corresponding to R2. Claude Opus 4.7 scored 92.60 this period, up 14.9 points from the previous period, and has strong recovery capability in R3; it can be considered for scenarios requiring rapid recovery after a breach, but its S_integrity honest self-reporting dimension still needs monitoring.

Strategic Judgment: Signals That Adherence Is Underestimated and Overestimated

Based on this round's data, DeepSeek V4 Pro's adherence capability may be underestimated by the market: its 97.70 score is only 1 point behind the second-place model, yet it maintains the highest stability under the worst-of-3 standard. GLM-4.6's 76.50 and Qwen3 Max's 75.20 are close, and both fail at R1, suggesting that the two models share a similar mechanism defect in the constraint injection stage; whether this is a common training-stage weakness deserves focused testing next period. Gemini 3.1 Pro improved by 17.4 points from the previous period and achieved full marks across R1-R3, showing that its recovery mechanism under v3 multi-turn gradual pressure is approaching the first tier.

Global statistics show a 60% full-score rate and an R3 crash rate of only 0.9%, indicating that most current models can still maintain constraints under single-turn high pressure, but the worst-of-3 standard has exposed the real risks of tail-end models. When selecting models, enterprises are advised to make models scoring above 90 on WDCD the default option for production workflows, use models scoring below 90 only in low-risk, non-persistent-state scenarios, and add external checks.

The gap between DeepSeek V4 Pro's 97.70 and Qwen3 Max's 75.20 is no longer just a score difference; it is the practical dividing line between whether a production system can maintain initial constraints across 8-12 rounds of dialogue.

Data: YZ Index WDCD Adherence Leaderboard | Run #326 · Overall Leaderboard Ranking | Evaluation Methodology

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!