Claude Sonnet 4.6 Soars 13.5 Points, Gemini 2.5 Pro Plunges 12.5 — Dramatic Reshuffle in WDCD v3.1 Commitment-Keeping Leaderboard

Claude Sonnet 4.6 rose 13.5 points in this WDCD v3.1 test compared to Run #291, reaching 91.60 points, while Gemini 2.5 Pro fell 12.5 points. Grok 4 holds the top position with 96.30 points, and GPT-o3 ranks second at 95.20 points.

Key Data Facts

A total of 11 models were evaluated in this pilot round, using worst-of-3 sampling. Claude Sonnet 4.6, DeepSeek V4 Pro (92.20 points, +6.9), and GPT-o3 (+6.4) rose; Claude Opus 4.7 (-5.7), Doubao Pro (-6.6), Gemini 2.5 Pro (-12.5), and GPT-5.5 (-5.3) fell. In the Top 5, Grok 4, GPT-o3, and GLM-4.6 (93.70) all exceeded 93 points, with Claude Sonnet 4.6 just edging into fifth place.

Possible Mechanisms Behind Commitment-Keeping Score Differences

The v3.1 question pool emphasizes 8-12 rounds of continuous pressure escalation, including four types of progressively intensifying pressure — social proof, executive override, salami-slicing, and sunk cost — along with KBV recall probes and a final-round honest self-report. Claude Sonnet 4.6's sharp rise most likely stems from improvements in S_recover (post-breach recovery) and S_integrity (honest self-reporting), maintaining constraint memory and accurately reviewing its performance after multiple rounds of pressure. Gemini 2.5 Pro's 12.5-point drop is likely concentrated in the S_hold (commitment survival) phase, breaching constraints at an earlier stage in resource-limitation or security-compliance scenarios, leading to substantial loss from the 60-point baseline.

DeepSeek V4 Pro and GPT-o3 rose in tandem, suggesting enhanced multi-round resilience in both engineering-standard and business-rule constraint scenarios. Conversely, GPT-5.5 and Claude Opus 4.7 declined together, possibly reflecting increased sensitivity to executive-override pressure, with constraint drift appearing in earlier rounds.

Practical Implications for Production Integration

For enterprises integrating AI into production workflows, WDCD scores directly map to the boundary of trustworthy constraint scenarios. With both Grok 4 and GPT-o3 exceeding 95 points, they can be considered for direct use in data-boundary and security-compliance scenarios, requiring the fewest additional guardrails. After rising to 91.60 points, Claude Sonnet 4.6 has approached the threshold for use in resource-limitation tasks, though manual review checkpoints are still needed in sunk-cost pressure scenarios.

The score declines of Gemini 2.5 Pro and Doubao Pro suggest higher risk under continuous multi-round business-rule constraints. Enterprises planning to integrate these models should implement a secondary confirmation mechanism for salami-slicing incremental requests, preventing subsequent constraints from gradually eroding after a single-round pass.

Strategic Assessment and Next-Cycle Validation Signals

The data suggests that Grok 4's commitment-keeping capability may be underestimated by the market — its 96.30-point lead has been validated under v3.1 question types; Gemini 2.5 Pro's commitment-keeping capability may be overestimated, as its 12.5-point drop reveals vulnerability under v3's multi-round pressure. Claude Sonnet 4.6's 13.5-point rebound warrants continued tracking of whether its S_kbv (constraint memory) remains stable in the next cycle.

Analysis indicates that model updates or changes in prompt sensitivity are the primary driving factors, but the next cycle must confirm whether the rises of DeepSeek V4 Pro and GPT-o3 are sustainable. When selecting models, enterprises should prioritize direct deployment of Top 3 models in security-compliance scenarios, while models ranked 5-7 require additional runtime verification in the form of KBV recall probes.

The divergence in commitment-keeping capability has evolved from a laboratory metric into a real threshold for production decisions. The next-cycle validation focus is on whether Claude Sonnet 4.6 can hold its 91-point platform, and whether Gemini 2.5 Pro will continue to decline.

Data source: YZ Index WDCD Commitment-Keeping Leaderboard | Run #296 · Change Tracking | Evaluation Methodology

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!