In the WDCD v3.1 commitment-adherence test, Grok 4 became the only full-score model with 100.00 points, while Doubao Pro ranked 11th with 75.30 points, a score gap of 24.7 points between the two models.
Ranking Landscape: Concentrated Top, Fractured Bottom
The score distribution of the 11 models in this round shows clear tiers. Grok 4 (100.00), GLM-4.6 (96.40), and Gemini 3.1 Pro (94.90) form the first tier, all above 94 points. Claude Opus 4.7 (93.30) and GPT-o3 (93.10) follow closely, with a gap of less than 1 point. After DeepSeek V4 Pro (91.40), Gemini 2.5 Pro (89.10), Claude Sonnet 4.6 (87.90), Qwen3 Max (87.50), and GPT-5.5 (86.40) form the second tier. There is still a 12.2-point gap between Doubao Pro (75.30) and ninth-place Qwen3 Max.
The overall full-score rate was 64.5%, and the R3 breakdown rate was 0%. This indicates that after 8-12 rounds of continuous pressure on v3 questions, all models avoided being completely broken in the final round, but the score differences mainly came from S_hold commitment survival and S_recover breakdown recovery.
Champion Analysis: How Grok 4 Achieved Zero Deductions
Grok 4 scored full marks on all 10 v3 questions and 8 v2 anchor questions, with no deductions in any of the four components: S_hold, S_kbv, S_recover, and S_integrity. Under worst-of-3 sampling, its worst run still held 100 points, indicating that across the five constraint scenarios—data boundaries, resource limits, business rules, safety compliance, and engineering specifications—Grok 4 fully retained its established constraints through the KBV restatement probe and the final-round honest debriefing.
By contrast, second-place GLM-4.6 scored 96.40, with deductions mainly concentrated in S_recover. The 3.6-point difference between Grok 4 and GLM-4.6 reflects that Grok 4's constraint memory was more stable during consecutive rounds of social proof and sunk-cost pressure.
Bottom Analysis: The Mechanism Weaknesses Behind Doubao Pro's 75.30
On v3 questions, Doubao Pro's S_hold commitment survival score was significantly lower than those of the top models, leaving its total at only 75.30. After conversion for the v2 anchor questions, its R3 pressure-phase score was a clear drag. Although the R3 breakdown rate was 0% across the board, after multiple rounds of salami-slicing pressure, Doubao Pro's completeness in retaining constraints was insufficient, and it also had deductions in S_integrity honest self-reporting.
Compared with ninth-place Qwen3 Max (87.50), Doubao Pro was 12.2 points lower, with the main gap appearing in 8-12 turn conversations in the two constraint scenarios of safety compliance and engineering specifications.
Causes of the Gap Between Top and Bottom Tiers
The score differences most likely stem from the two components S_hold and S_recover. After the commitment-establishment phase (2-5 parallel hard constraints), top models could maintain constraint memory under escalating pressure from authority overrides and sunk costs, while bottom models already showed constraint drift at the KBV restatement probe stage.
GLM-4.6 improved by 19.9 points this period and Qwen3 Max by 20.2 points, indicating clear progress in their S_recover capability on v3 multi-turn gradual-pressure questions. Grok 4 improved by only 6.4 points but still maintained a full score, showing that its baseline commitment-adherence capability is already at a high level.
Selection Implications for Production Workflow Integration
Enterprises integrating AI into production workflows may prioritize models scoring above 95 on WDCD for safety compliance and data boundary tasks. Grok 4, GLM-4.6, and Gemini 3.1 Pro still maintain high scores in their worst-of-3 performance, making them suitable for direct integration in scenarios that do not require additional multi-turn review.
For resource-limit and business-rule tasks, it is advisable to add independent guardrails for models scoring below 90, for example by mandating a KBV restatement probe after the R2 interference round. For a model like Doubao Pro at 75.30, additional manual spot checks are needed in engineering-specification scenarios to prevent premature commitment breach during the S_hold phase.
Strategic Assessment
Grok 4's commitment-adherence capability may be underestimated by the market; its full score of 100 and 0 breakdown rate indicate that it has formed a technical moat under the current v3.1 question pool. The 19.9-point and 20.2-point gains by GLM-4.6 and Qwen3 Max show that open-source/domestic models are rapidly catching up in recovery capability under continuous pressure.
Signals worth verifying next period are whether GLM-4.6 can further narrow its 3.6-point gap with Grok 4 on S_hold, and whether Doubao Pro can raise its score above 80 in the R3 pressure phase. Current data only support the conclusion that "Grok 4 is the most stable performer across the five constraint scenarios."
A full-score model has emerged. The gap is not accidental, but an inevitable divergence in constraint memory under multi-turn pressure.
Data source: YZ Index WDCD Commitment-Adherence Leaderboard | Run #346 · Overall Leaderboard Ranking | Evaluation Methodology
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接