Grok 4 Tops WDCD Commitment Ranking with 93.80 Points; GLM-4.6 Ranks Last with 68, a 25.8-Point Gap

The WDCD v3.1 commitment test shows Grok 4 ranking first with 93.80 points and GLM-4.6 ranking last with 68.00 points, a difference of 25.8 points between them. The general performance of the 11 evaluated models shows a clear divide: the top four models all exceed 83 points, while the bottom three fall below 72 points.

Ranking Structure and Divergence Under R3 Pressure

This test uses worst-of-3 sampling, with v3 questions and v2 anchor questions averaged on equal weights. Grok 4 achieves R1=1.00, R2=1.00, and R3=1.50/2 on v2 anchor questions, retaining 75% of its R3 score after three rounds of progressive interference. GPT-o3 follows closely with 89.80 points, but its R3 only scores 0.50/2, 1 point lower than Grok 4. Claude Sonnet 4.6 ranks third with 87.20 points; it earns full marks of 2.00/2 on R3 yet scores 0 in the R2 phase, indicating notable constraint loosening by the second round of continuous pressure.

The common feature of the tail-end models is a complete collapse in the R3 phase. GLM-4.6 scores 0 in R1, R2, and R3, showing that it cannot maintain any anchor after the initial constraint injection. Qwen3 Max scores 1 point in both R1 and R2 but 0 in R3, finishing with only 71.60 points. Gemini 2.5 Pro and Claude Opus 4.7 likewise drop to zero in the R3 phase, ranking tenth and eighth, respectively.

Cause Analysis: R3 Collapse and Constraint Scenario Distribution

The overall R3 collapse rate stands at 9.1%, but collapses among tail-end models are concentrated in two types of constraint scenarios: resource limits and security compliance. GLM-4.6 fails to withstand R3 pressure across all five scenario categories, suggesting it is most likely to abandon initial hard constraints during the sunk-cost pressure-escalation phase. Qwen3 Max performs reasonably in the data-boundary scenario but breaks down outright in the R3 phase of the engineering-specification scenario, indicating that its constraint memory is scenario-dependent.

Differences among the top models mainly stem from the S_hold commitment-survival metric. Grok 4's high score comes from a later point of breach rather than from never breaching at all. Although Claude Sonnet 4.6 achieves a perfect R3 score, its S_kbv constraint-memory score is low, showing that partial forgetting occurs during the KBV recitation probe phase.

Implications for Model Selection in Production Workflow Integration

Enterprises integrating AI into production workflows should pay close attention to R3-phase performance. Grok 4 retains 1.50 points on R3 in the resource-limit and business-rule scenarios, making it suitable for use cases that require multi-round tool calls and budget-cap controls. GPT-o3 and Claude Sonnet 4.6 diverge sharply in R3 performance within the security-compliance scenario: the former scores only 0.50 points on R3 while the latter earns full marks. Approval workflows that touch compliance boundaries are therefore better served by Claude Sonnet 4.6.

Tail-end models should not be directly deployed in pipelines that require long-term constraint maintenance. GLM-4.6 scores 0 on R3 across all scenarios, meaning that once a conversation enters sustained pressure, the initially defined data boundaries or engineering specifications can easily be breached. If enterprises must use such models, they should add external guardrails at the R2 phase rather than relying on the model's own commitment-keeping ability.

Strategic Assessment and Next-Round Verification Signals

Claude Opus 4.7 dropped 17.0 points this round compared with the previous round and scored 0 in the R2 phase, showing clear regression in its ability to maintain constraints under continuous interference — possibly linked to the added salami-slicing pressure rounds in the v3 question pool. Both Gemini models fell by more than 18 points, with Gemini 2.5 Pro dropping to tenth place, suggesting a notable disconnect between their performance under v3's multi-round progressive pressure and their results on v2 anchor questions.

DeepSeek V4 Pro and GPT-o3 each dropped 5.0 points, but the former still retains 1.00 point on R3 while the latter has only 0.50, indicating that DeepSeek may hold an edge in recovery capability. Qwen3 Max is the only model that scores full marks on R2 yet drops to zero on R3, and its performance under social-conformity pressure merits focused tracking in the next round.

The worst-of-3 basis used in this pilot phase amplified the volatility of tail-end models; if scoring shifts to an average-based approach in the future, the relative positions of GLM-4.6 and Gemini 2.5 Pro may change. When selecting models, enterprises should prioritize those with R3 scores above 1.00 to reduce the risk of constraint failure in production environments.

Commitment-keeping ability is not a byproduct of model parameters; it is a hard threshold that determines whether a model can truly enter the production loop.

Data source: YZ Index WDCD Commitment Ranking | Run #316 · Overall Leaderboard | Evaluation Methodology

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!