Gemini 2.5 Pro's WDCD Score Surges 25.1 Points; Six Models Improve, Only Doubao Pro Drops 6.3

In this round of WDCD v3.1 testing, Gemini 2.5 Pro rose 25.1 points in a single period, Gemini 3.1 Pro rose 17.4 points, DeepSeek V4 Pro became the current highest-scoring model with 97.70 points, six models showed positive changes, and only Doubao Pro declined by 6.3 points.

Data Facts: The Precise Distribution of Gains and Declines

Compared with Run #316, among the 11 evaluated models this period, Claude Opus 4.7 rose 14.9 points, DeepSeek V4 Pro rose 14.1 points, GLM-4.6 rose 8.5 points, and GPT-5.5 rose 9.9 points. The Top 5 list shows DeepSeek V4 Pro at 97.70 points, Gemini 3.1 Pro at 96.70, Grok 4 at 96.00, Gemini 2.5 Pro at 94.50, and GPT-o3 at 94.20. The sampling uses a worst-of-3 criterion, with the total score derived by averaging v3 questions and v2 anchor questions at equal weight.

Cause Analysis: Mechanism Differences Under Multi-Round Pressure

WDCD v3 questions use 8-12 rounds of dialogue: first establish 2-5 hard constraints, then sequentially apply social proof, authority special approval, salami-slicing, and sunk-cost pressure, and finally a KBV restatement probe and a final-round honest self-report. The large gains by Gemini 2.5 Pro and Gemini 3.1 Pro most likely came from simultaneous improvement in the 60-point S_hold commitment-survival score and the 10-point S_recover breach-recovery score, indicating that they break commitments later and recover faster under consecutive pressure rounds. DeepSeek V4 Pro's total score of 97.70 points suggests standout performance in the 15-point S_kbv constraint-memory score in data-boundary and safety-compliance scenarios. Doubao Pro's 6.3-point decline may stem from deductions on the 15-point S_integrity honest self-report, meaning it failed to self-report accurately after breaking a commitment, reflecting a weakness in the final-round review stage. The scaled portion of the three-round v2 anchor questions, with a maximum of 4 points, also shows that rising models scored higher in the R3 pressure round.

Selection Implications: Practical Risk Boundaries for Production Integration

For enterprises integrating AI into production workflows, WDCD scores directly correspond to the strength of constraints they can bear. DeepSeek V4 Pro and Gemini 3.1 Pro both score above 96, so they can be used directly in resource-limited and business-rule scenarios, with less need for additional guardrails. Although Gemini 2.5 Pro scores 94.50, its single-period gain of 25.1 points makes it suitable for piloting first in low-risk engineering-standards tasks. Doubao Pro did not enter the Top 5, so enterprises should add an independent validation layer in safety-compliance scenarios and avoid relying on its S_recover capability. All models still require human review in data-boundary scenarios, because the worst-of-3 criterion has already exposed the worst single performance.

Strategic Judgment: Underestimated and Overestimated Signals

DeepSeek V4 Pro currently leads with 97.70 points and, together with its 14.1-point gain, suggests the market had previously underestimated its commitment-adherence capability. The simultaneous large gains by the two Gemini models indicate a systemic improvement in their underlying robustness to multi-round incremental pressure. Doubao Pro's 6.3-point decline is the only negative signal this period and warrants focused examination next period of its specific performance in the sunk-cost pressure round. Although Grok 4 and GPT-o3 did not have change data listed, they remain firmly in the Top 5, showing that their baseline commitment-adherence capability remains competitive. In the next cycle, close attention should be paid to whether Claude Opus 4.7 and GLM-4.6 can convert their respective 14.9-point and 8.5-point gains into Top 5 spots.

Commitment adherence is no longer a static label but a real-time survival contest after each round of pressure; DeepSeek V4 Pro and the Gemini series have used data to prove who can push the moment of commitment breach later.

Data from: YZ Index WDCD Commitment-Adherence Leaderboard | Run #326 · Change Tracking | Evaluation Methodology

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!