GPT-o3 tops the WDCD v3.1 Commitment Ranking with 94.00 points, while Qwen3 Max ranks last with 62.60 points, showing a 31.4-point gap between the leading and trailing models among 11 evaluated models.
Ranking Structure: Three-Tier Divergence and the Key Role of R3
This WDCD evaluation uses worst-of-3 sampling, with equal weighting of v3 questions and v2 anchor questions. GPT-o3, Grok 4, and Claude Opus 4.7 form the first tier, scoring 94.00, 87.90, and 87.60 respectively, with gaps of less than 7 points between them. The second tier ranges from Gemini 3.1 Pro at 87.30 to GLM-4.6 at 78.30, a 9-point spread mainly driven by differences in R3 stress scores. At the bottom, Doubao Pro at 63.40 and Qwen3 Max at 62.60 are clearly distanced from the mid-range.
R3 has a collapse rate of only 3.6%, yet it directly determines the final ranking. DeepSeek V4 Pro scored full 2 points in R3 but still ranks fifth overall, indicating significant losses in the v3 multi-round gradual pressure stage. In contrast, Gemini 3.1 Pro scored 0 in R3 but maintained fourth place thanks to high scores in v3's S_hold and S_kbv, proving that early constraint memory can partially compensate for a final-round collapse.
Root Cause Analysis: Constraint Scenarios and Pressure Round Differences
Among the five constraint scenarios, safety compliance and engineering specification scenarios are most prone to triggering R3 collapse. GPT-o3 maintains R1=1, R2=1, R3=1 across all scenarios, showing the strongest resistance to the gradual pressure escalation of "salami slicing" and "sunk costs." Qwen3 Max and GLM-4.6 repeatedly scored 0 in the R3 phase, suggesting they are more likely to abandon initial constraints under social proof pressure in resource limitation and business rule scenarios.
Claude Sonnet 4.6 improved by 15.0 points this round, mainly due to recoveries in v3's S_recover and S_integrity, indicating significant progress in honest self-reporting ability after KBV repetition probes. GLM-4.6, on the other hand, dropped 15.3 points; combined with its record of scoring 0 in R3, this points to faster decay of constraint memory over 8-12 consecutive conversation rounds.
Selection Implications: Practical Risk Boundaries for Production Integration
For enterprises integrating AI into production workflows, GPT-o3 with a WDCD score of 94.00 can be used directly in data boundary and safety compliance scenarios, with guardrails minimized. Grok 4 and Claude Opus 4.7, both in the 87-point range, are suitable for resource-constrained tasks, though manual review checkpoints should remain in engineering specification scenarios.
Qwen3 Max and Doubao Pro, scoring 62-63, show significantly higher breach probabilities in business rule scenarios. Enterprises should add an independent constraint validation layer before these models, or use them only for low-risk internal Q&A. DeepSeek V4 Pro, despite a perfect R3 score, ranks only 84.30 overall, which is still below the first tier; it is recommended to overlay external memory anchors for multi-round long conversations.
Strategic Judgment: Signals Underestimated and Overestimated
Claude Sonnet 4.6's single-round improvement of 15.0 points suggests its commitment ability may have been underestimated by the market. If it maintains its R3 score in the next round, it could enter first-tier competition. GLM-4.6's data over two consecutive rounds shows vulnerability in the R3 phase, warranting focused verification of its recovery path in safety compliance scenarios.
GPT-o3's leading score of 94.00 is built on maintaining high scores even in the worst-of-3 worst-case outcome, indicating a replicable mechanistic advantage in commitment stability. The 31.4-point gap between trailing and leading models is unlikely to be closed quickly through single-version iterations. Enterprises should treat WDCD scores as a hard entry requirement, not a reference indicator, when selecting models.
Commitment ability is not a byproduct of model size, but a core moat for production usability.
Data source: YZ Index WDCD Commitment Ranking | Run #233 · Overall ranking | Evaluation Methodology
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接