In the WDCD v3.1 commitment test, Gemini 3.1 Pro ranked first among the 11 evaluated models with 97.70 points, while Qwen3 Max ranked last with 70.30 points, a gap of 27.4 points. Sampling was conducted under the worst-of-3 methodology, with the question pool comprising 17 v3 multi-turn progressive-pressure questions and 8 v2 anchor questions. Scoring strictly followed the four-component framework: S_hold (60 points), S_kbv (15 points), S_recover (10 points), and S_integrity (15 points).
Ranking Landscape: Concentrated at the Top, Cliff-Like Drop at the Bottom
The top five models all scored above 94.7 on WDCD: Gemini 3.1 Pro 97.70, Grok 4 96.30, GLM-4.6 95.00, GPT-o3 94.80, and Claude Opus 4.7 94.70. Sixth and seventh place, Gemini 2.5 Pro at 93.20 and GPT-5.5 at 92.60, remained in the above-90 range. The decline began with eighth-place DeepSeek V4 Pro at 88.60, followed by a pronounced cliff: ninth-place Doubao Pro at 79.10, tenth-place Claude Sonnet 4.6 at 78.60, and eleventh-place Qwen3 Max at 70.30.
Global statistics show a full-score rate of 68.2% and an R3 collapse rate of only 2.7%. This indicates that most models can maintain constraints during the R1 and R2 stages of v2 anchor questions, but differences widen in the R3 pressure stage. GLM-4.6 scored 0.00/2 in R3, GPT-o3 scored 0.00 in R2, DeepSeek V4 Pro scored 0.00 in both R2 and R3, and Qwen3 Max scored 0.00 in R3, showing that tail-end models see significantly lower S_hold scores under sustained pressure.
Champion Analysis: Gemini 3.1 Pro's Constraint Memory and Recovery Mechanism
Gemini 3.1 Pro achieved perfect scores on v2 anchor questions (R1=1.00, R2=1.00, R3=2.00/2). Combined with its high S_kbv and S_recover scores on v3 questions, this suggests it can effectively resist social-conformity and sunk-cost pressure in data-boundary and safety-compliance scenarios. The likely cause is its more stable responses to KBV reiteration probes across multiple parallel hard constraints, along with higher final-round honest self-report scores.
For enterprises running production workflows, this model can be integrated directly in resource-constraint and engineering-standard scenarios, with strong recovery capability after a breach and no need for additional guardrails. However, manual review is still recommended in business-rule scenarios, because under the S_integrity 15-point mechanism, any false claim of a clean record results in an immediate score of 0.
Last-Place Analysis: Qwen3 Max's R3 Collapse and S_hold Loss
Qwen3 Max scored 70.30 on WDCD, with v2 anchor question results of R1=1.00, R2=1.00, and R3=0.00/2. The complete loss of R3 full marks directly caused its significant lag after total score aggregation. Combined with worst-of-3 sampling on v3 questions, its largest losses occurred in the S_hold 60-point component during consecutive pressure rounds, suggesting that its constraint memory decays more rapidly under authoritative special-approval and salami-slicing pressure.
The model-selection implication is clear: in safety-compliance and data-boundary scenarios, this model should not be used directly in production workflows and requires additional output-filtering guardrails. Otherwise, once a multi-turn conversation leads to a breach, the probability of recovery is low, and S_recover scores cannot compensate for the loss.
Mechanistic Differences Behind the Leading Tier vs. the Trailing Gap
Leading models (top five) averaged above 1.6/2 in the R3 stage, while trailing models (bottom three) averaged 0.67/2. The gap is concentrated in the continuous-pressure stage of v3 questions and the R3 pressure stage of v2 anchor questions. GLM-4.6 and Qwen3 Max both scored 0 in R3, indicating their S_hold survival time is short under escalating sunk-cost pressure.
Strategic assessment: Claude Opus 4.7 scored 94.70 this round with a perfect R3, and its commitment-keeping capability may be undervalued by the market. Claude Sonnet 4.6 scored 78.60, a 13-point gap within the same product family, with unstable R2 and R3 performance, a signal worth validating in the next round. DeepSeek V4 Pro scored 88.60 but recorded 0 in both R2 and R3, posing higher integration risk in engineering-standard scenarios.
Round-over-round comparisons support the following observations: Gemini 2.5 Pro rose 22.2 points and Doubao Pro rose 11.2 points, indicating improved S_recover and S_integrity scores on v3 questions; Claude Sonnet 4.6 fell 13.0 points, suggesting possible constraint-memory issues during the R2 stage.
For enterprise model selection, models scoring above 90 are suitable for direct integration in safety-compliance and data-boundary scenarios; models in the 80-90 range require added guardrails in resource-constraint scenarios; and models below 70 are recommended only for low-risk business-rule testing. All of these assessments are derived from this round's WDCD scores and R1-R3 component-level data.
Commitment-keeping capability is not an accessory of model parameters; it is the core moat determining whether production workflows can remain stable over the long term.
Data source: YZ Index WDCD Commitment Leaderboard | Run #306 · Overall Leaderboard | Evaluation Methodology
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接