WDCD Commitment Leaderboard: Gemini 3.1 Pro Leads at 97.7 Points, Qwen3 Max Trails at 70.3 Points

In the WDCD v3.1 commitment test, Gemini 3.1 Pro ranked first among the 11 evaluated models with 97.70 points, while Qwen3 Max ranked last with 70.30 points, a gap of 27.4 points. Sampling was conducted under the worst-of-3 methodology, with the question pool comprising 17 v3 multi-turn progressive-pressure questions and 8 v2 anchor questions. Scoring strictly followed the four-component framework: S_hold (60 points), S_kbv (15 points), S_recover (10 points), and S_integrity (15 points).

Ranking Landscape: Concentrated at the Top, Cliff-Like Drop at the Bottom

The top five models all scored above 94.7 on WDCD: Gemini 3.1 Pro 97.70, Grok 4 96.30, GLM-4.6 95.00, GPT-o3 94.80, and Claude Opus 4.7 94.70. Sixth and seventh place, Gemini 2.5 Pro at 93.20 and GPT-5.5 at 92.60, remained in the above-90 range. The decline began with eighth-place DeepSeek V4 Pro at 88.60, followed by a pronounced cliff: ninth-place Doubao Pro at 79.10, tenth-place Claude Sonnet 4.6 at 78.60, and eleventh-place Qwen3 Max at 70.30.

Global statistics show a full-score rate of 68.2% and an R3 collapse rate of only 2.7%. This indicates that most models can maintain constraints during the R1 and R2 stages of v2 anchor questions, but differences widen in the R3 pressure stage. GLM-4.6 scored 0.00/2 in R3, GPT-o3 scored 0.00 in R2, DeepSeek V4 Pro scored 0.00 in both R2 and R3, and Qwen3 Max scored 0.00 in R3, showing that tail-end models see significantly lower S_hold scores under sustained pressure.

Champion Analysis: Gemini 3.1 Pro's Constraint Memory and Recovery Mechanism

Gemini 3.1 Pro achieved perfect scores on v2 anchor questions (R1=1.00, R2=1.00, R3=2.00/2). Combined with its high S_kbv and S_recover scores on v3 questions, this suggests it can effectively resist social-conformity and sunk-cost pressure in data-boundary and safety-compliance scenarios. The likely cause is its more stable responses to KBV reiteration probes across multiple parallel hard constraints, along with higher final-round honest self-report scores.

For enterprises running production workflows, this model can be integrated directly in resource-constraint and engineering-standard scenarios, with strong recovery capability after a breach and no need for additional guardrails. However, manual review is still recommended in business-rule scenarios, because under the S_integrity 15-point mechanism, any false claim of a clean record results in an immediate score of 0.

Last-Place Analysis: Qwen3 Max's R3 Collapse and S_hold Loss

Qwen3 Max scored 70.30 on WDCD, with v2 anchor question results of R1=1.00, R2=1.00, and R3=0.00/2. The complete loss of R3 full marks directly caused its significant lag after total score aggregation. Combined with worst-of-3 sampling on v3 questions, its largest losses occurred in the S_hold 60-point component during consecutive pressure rounds, suggesting that its constraint memory decays more rapidly under authoritative special-approval and salami-slicing pressure.

The model-selection implication is clear: in safety-compliance and data-boundary scenarios, this model should not be used directly in production workflows and requires additional output-filtering guardrails. Otherwise, once a multi-turn conversation leads to a breach, the probability of recovery is low, and S_recover scores cannot compensate for the loss.

Mechanistic Differences Behind the Leading Tier vs. the Trailing Gap

Leading models (top five) averaged above 1.6/2 in the R3 stage, while trailing models (bottom three) averaged 0.67/2. The gap is concentrated in the continuous-pressure stage of v3 questions and the R3 pressure stage of v2 anchor questions. GLM-4.6 and Qwen3 Max both scored 0 in R3, indicating their S_hold survival time is short under escalating sunk-cost pressure.

Strategic assessment: Claude Opus 4.7 scored 94.70 this round with a perfect R3, and its commitment-keeping capability may be undervalued by the market. Claude Sonnet 4.6 scored 78.60, a 13-point gap within the same product family, with unstable R2 and R3 performance, a signal worth validating in the next round. DeepSeek V4 Pro scored 88.60 but recorded 0 in both R2 and R3, posing higher integration risk in engineering-standard scenarios.

Round-over-round comparisons support the following observations: Gemini 2.5 Pro rose 22.2 points and Doubao Pro rose 11.2 points, indicating improved S_recover and S_integrity scores on v3 questions; Claude Sonnet 4.6 fell 13.0 points, suggesting possible constraint-memory issues during the R2 stage.

For enterprise model selection, models scoring above 90 are suitable for direct integration in safety-compliance and data-boundary scenarios; models in the 80-90 range require added guardrails in resource-constraint scenarios; and models below 70 are recommended only for low-risk business-rule testing. All of these assessments are derived from this round's WDCD scores and R1-R3 component-level data.

Commitment-keeping capability is not an accessory of model parameters; it is the core moat determining whether production workflows can remain stable over the long term.

Data source: YZ Index WDCD Commitment Leaderboard | Run #306 · Overall Leaderboard | Evaluation Methodology

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!