Grok 4 Leads WDCD Commitment-Keeping Leaderboard with 91.76 as GLM-4.6 Trails at 61.55

Grok 4 ranks first on the WDCD Commitment-Keeping Leaderboard with a WDCD score of 91.76, while GLM-4.6 comes last with 61.55, leaving a 30.21-point gap between the top and bottom.

Ranking landscape: concentrated at the top, a cliff at the bottom

This WDCD v3.1 pilot test covered 11 models and used worst-of-3 sampling. Grok 4, Gemini 3.1 Pro, and GPT-o3 formed the first tier, scoring 91.76, 89.59, and 87.86, respectively. Fourth-place Claude Opus 4.7 scored 83.97, 7.79 points behind first place, while the gap within the tier was less than 4 points.

Seventh-place Doubao Pro at 79.69 to tenth-place GPT-5.5 at 73.76 formed the second tier, spanning 5.93 points. GLM-4.6 ranked alone at the bottom with 61.55, 12.21 points below tenth place and 30.21 points below first place.

Champion analysis: Grok 4 shows the highest resilience in the R3 pressure phase

Grok 4's v2 anchor-question scores were R1=1.00, R2=0.88, and R3=1.25/2, which after conversion supported its overall lead. The R3 phase corresponds to sunk cost and salami-slicing tactics under continuous pressure; Grok 4 scored highest in this round, indicating the strongest ability to keep commitments and survive under multi-turn, gradual pressure.

By contrast, Gemini 3.1 Pro scored slightly higher at 1.38/2 in R3, but with the same R2 score, it still trailed Grok 4 by 2.17 points overall. GPT-o3 scored only 0.75/2 in R3, dragging its total down to third place.

Bottom analysis: GLM-4.6 completely fails all anchor questions

GLM-4.6 scored 0.00/2 in R1, R2, and R3, down 29.8 points from the previous period. This is the only model in this test that scored zero across all three rounds of v2 anchor questions, indicating that it failed to maintain its initial commitments across the constraint-injection, interference, and pressure phases.

The model's S_hold commitment-survival score on v3 questions may be extremely low, making it difficult to offset the overall score even with the 15-point S_integrity component. The global R3 collapse rate was 9.4%, to which GLM-4.6 contributed a significant portion.

Mechanisms behind the gap between the top tier and the bottom

The gap is concentrated mainly in the continuous-pressure phase of v3 questions and the R3 round of v2 anchor questions. Over 8-12 dialogue turns, leading models still maintained relatively high S_hold scores for the initial 2-5 hard constraints when facing social proof, special authority approval, and salami-slicing tactics. Trailing models had already broken before the KBV paraphrase probe, and their S_recover capability was insufficient.

Differences are most easily exposed in data security boundaries and safety compliance scenarios. GLM-4.6 may fail earliest under engineering-specification constraints, whereas Grok 4 performs more stably in resource-limit and business-rule scenarios.

Selection implications for production workflow integration

For enterprises integrating AI into production workflows, choosing models scoring above 90 on WDCD in safety compliance and data-boundary scenarios can reduce guardrail investment. The S_hold scores of Grok 4 and Gemini 3.1 Pro support less real-time monitoring.

In resource-limit and engineering-specification scenarios, models below 73 require additional post-hoc audits and manual review. GLM-4.6's 61.55 indicates it is unsuitable for deployment without guardrails, and any business rule involving hard constraints requires an additional isolation layer.

For business-rule scenarios, Claude Opus 4.7 (83.97) can be considered as a balanced option, with an R3 score of 1.25/2 close to Grok 4.

Strategic judgment: underestimated commitment-keeping ability and signals to verify

Grok 4's commitment-keeping ability may be underestimated by the market. Its R3 score of 1.25/2 at 91.76 overall shows that it maintains relatively high integrity even under the most severe pressure, and its S_recover performance under multi-turn gradual pressure deserves focused verification in the next v3.2 period.

GLM-4.6's 29.8-point decline signals a clear regression in constraint memory. The next period should observe whether it can recover a non-zero score in the R1 phase of v2 anchor questions.

GPT-5.5's 73.76 is down 6.0 points from the previous period, with an R3 score of only 0.88/2, suggesting that its recovery ability under sunk-cost pressure may be overestimated.

Commitment-keeping ability is not an appendage of model scale, but a real moat for production deployment.

Data source: YZ Index WDCD Commitment-Keeping Leaderboard | Run #331 · Overall Ranking | Evaluation Methodology

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!