Grok 4 leads this commitment-keeping test with a WDCD score of 93.21, while Qwen3 Max ranks 11th with 74.31 — a spread of 18.9 points between the two.
Ranking Landscape and Key Metrics
A total of 11 models took part in this WDCD v3.1 test, using worst-of-3 sampling on a pool of 29 questions. Grok 4 posted the highest score after conversion across the v3 multi-round progressive-pressure items and v2 anchor items, with R1=1.00, R2=1.00, and R3=1.25/2. GLM-4.6 follows with 91.38 points, reaching R3 of 1.50/2. DeepSeek V4 Pro (87.97), GPT-o3 (86.93), and Gemini 3.1 Pro (85.72) form the second tier. At the tail, Qwen3 Max has an R3 of only 0.63/2; overall, the R3 collapse rate is 9.1% and the full-score rate is 58%.
Why the Champion Won: Constraint Survival in the R3 Pressure Stage
Across rounds 8–12 of the v3 items, the hard constraints Grok 4 established during the commitment phase were subjected to stepwise escalation through social approval, special authority approval, salami-slicing, and sunk costs. It earned the highest S_hold commitment-keeping survival score. The R3 pressure rounds directly determine a 60-point weighting, and Grok 4's R3 of 1.25/2 outperforms Claude Opus 4.7's 0.75/2 and GPT-5.5's 0.88/2. This indicates that under sustained pressure in engineering-norm and safety-compliance scenarios, Grok 4 holds out longer before breaking, while its S_recover recovery capability and S_integrity honest self-reporting also remain at relatively high levels.
Mechanism Shortfalls of the Bottom-Ranked Model
Qwen3 Max's R3 stands at only 0.63/2, with clear losses on the KBV restatement probe and the final-round honest review. Across the 11 models, the 9.1% R3 collapse rate is concentrated in the bottom three. Doubao Pro posted R3 of 1.13/2 yet only 75.62 overall, suggesting its S_kbv constraint-memory component (15 points) on v3 items may be weak. Claude Sonnet 4.6 fell 5.0 points from the previous round, with R2 dropping to 0.88 from a possibly higher level, showing weakened stability across the three-round anchor-item interference phase.
What the Gap Between Top and Bottom Means for Model Selection
For enterprises integrating AI into production workflows, Grok 4 and GLM-4.6's commitment-keeping data indicate they can be used directly in data-boundary and resource-constrained scenarios; safety-compliance scenarios should retain human review. Models with R3 below 1.00 — GPT-o3 (0.63/2) and Qwen3 Max (0.63/2) — are more likely to break under sustained business-rule pressure, so adding external guardrails or multi-model voting is recommended. Gemini 2.5 Pro has an R3 of 1.25/2 but totals only 77.48, indicating that the S_hold weighting on v3 items dragged down its overall performance; production deployment should include dedicated testing against sunk-cost escalation.
Strategic Takeaways: Underestimated Strength and Signals Worth Verifying
Judging by the correspondence between R3 scores and total scores, Grok 4's commitment-keeping capability may be underestimated by the market — its R3 of 1.25/2 is highly consistent with its 93.21 total. GLM-4.6, with an R3 of 1.50/2 but a total of 91.38, may be losing points in the S_integrity honest self-reporting component (15 points), which is worth verifying next round. Claude Opus 4.7 has an R2 of only 0.50, showing weaker constraint memory during the interference phase than its series stablemate Sonnet 4.6; strategically, it should not be the first choice for multi-round progressive-pressure scenarios. Doubao Pro rose 7.5 points from the previous round, and its R3 improvement to 1.13/2 may stem from v2 anchor-item optimization, though stability still needs to be observed.
Global statistics show that 58% of models achieved full scores in some constrained scenarios, indicating that the current v3.1 question pool already differentiates effectively among leading models. When selecting models, enterprises should prioritize those with R3 scores above 1.00 to reduce compliance risk in engineering-norm scenarios. Bottom-tier models score low on S_hold in safety-compliance and business-rule scenarios and require additional rule engines or manual review checkpoints.
Commitment-keeping is not a byproduct of model parameters; it is the invisible threshold for production deployment.
If the next test round adds more resource-constrained scenarios, R3 score differences may further widen the gap between top and bottom.
Data source: YZ Index WDCD Commitment-Keeping Leaderboard | Run #311 · Overall Ranking | Evaluation Methodology
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接