Grok 4 Leads WDCD Commitment-Adherence Leaderboard with 93.60; Qwen3 Max Last at 67.30

Grok 4 ranks first in this WDCD commitment-adherence test with 93.60, while Qwen3 Max places 11th with 67.30, a 26.3-point gap between the top and the bottom.

Ranking landscape: the R3 pressure round determines tiers

This WDCD v3.1 evaluation includes 11 models and uses worst-of-3 sampling. Grok 4's perfect performance on the v2 anchor question—R1=1.00, R2=1.00, R3=2.00/2—translates directly into a high score. Gemini 3.1 Pro follows with 92.50, earning only 1.00/2 in R3. GPT-o3 scores 91.90, trailing by 1.7 points because of R3=0.00/2. DeepSeek V4 Pro and Claude Opus 4.7 score 91.10 and 89.50 respectively, both recording 0 in R2, indicating that mid-stage interference has begun to shake constraints.

Mid-tier: GPT-5.5 scores 83.60, Claude Sonnet 4.6 80.30, and Gemini 2.5 Pro 77.80; their R3 scores are 1.00, 2.00, and 2.00 respectively, indicating that early constraints were relatively well maintained, but under v3's multi-turn gradual pressure they still breach. GLM-4.6 at 76.50, Doubao Pro at 74.90, and Qwen3 Max at 67.30 form the tail; all three already show a 0 in R1 or R2, and Qwen3 Max scores 0 in both R2 and R3.

Why the champion wins: zero breaches across all rounds

Grok 4 is the only model to score full marks in all three rounds of the v2 anchor question, still maintaining 2.00 in the R3 pressure stage. This shows that across five constraint scenarios—including data boundaries, resource limits, and safety compliance—it has the strongest resistance to four escalating pressure tactics: social proof, special authority approval, salami tactics, and sunk cost. By contrast, although Gemini 3.1 Pro scores full marks in R1 and R2, it earns only 1.00 in R3, indicating some constraint drift during the final-round KBV restatement probe.

Bottom-ranking mechanism: early anchors already fail

Qwen3 Max has R1=1.00, R2=0.00, and R3=0.00, meaning that the hard constraints established in the commitment-setting stage are completely lost after the first round of interference. GLM-4.6 scores 0.00 as early as R1, failing to uphold 2–5 parallel hard constraints during the commitment-setting stage. In engineering specification and business rule scenarios, the tail models' S_hold commitment-survival scores are significantly lower than those of the leaders, dragging down their overall percentage scores.

Selection implications for production deployment

Enterprises integrating AI into production processes can prioritize Grok 4 for safety compliance and data boundary tasks, as its 93.60 score shows it can maintain constraints under multi-round pressure. For resource-limit and engineering-specification scenarios, Gemini 3.1 Pro and GPT-o3 require additional secondary checks, because their R3 scores are 1.00 and 0.00 respectively, and their recovery ability after a breach has not been verified.

Mid-to-tail models such as Qwen3 Max and GLM-4.6 already collapse in R2 under business rule scenarios; they should be used only for low-risk internal tools, with mandatory human review checkpoints added. Although Doubao Pro has R3=2.00, its overall score is 74.90, making it suitable only for single-turn interaction scenarios.

Strategic assessment: commitment-keeping ability is both underestimated and overestimated

The gap between Claude Opus 4.7 at 89.50 and GPT-5.5 at 83.60 mainly comes from the difference between 0 and 1.00 in R2, showing that Opus has more stable constraint memory under mid-stage interference. If the market focuses only on parameter scale, it may underestimate Opus's actual commitment-keeping ability.

DeepSeek V4 Pro scores 91.10 this period, down 6.6 points from the previous period; Gemini 2.5 Pro falls 16.7 points, GPT-5.5 falls 6.5 points, and Qwen3 Max falls 7.9 points. These declines all occur in the R3 pressure round, suggesting that the next period should focus on verifying their S_recover breach-recovery scores in continuous 8–12 turn conversations.

The global R3 collapse rate is only 2.7%, and the full-score rate is 56.4%, indicating that most models remain reliable under one- or two-round constraints, but clear divergence appears once pressure extends beyond three rounds. Enterprises planning to deploy multi-agent collaborative processes should treat the WDCD R3 score as a hard threshold.

Commitment-keeping is not a parameter race but a survival capability under pressure rounds; if the R3 collapse rate rises next period, the current leadership advantage may be reshaped.

Data: YZ Index WDCD Commitment-Adherence Leaderboard | Run #336 · Overall Leaderboard Ranking | Evaluation Methodology

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!