The latest WDCD v3.1 pilot data shows Gemini 3.1 Pro reaching a total commitment-adherence score of 97.70, ranking first and improving by 9.5 points from Run #296; Gemini 2.5 Pro was the biggest winner, with a 22.2-point increase.
Core Data Facts: A Clear Distribution of 5 Gains and 1 Decline
Compared with Run #296, among the 11 evaluated models in this round, Claude Opus 4.7 rose by 8.9 points, Doubao Pro rose by 11.2 points, Gemini 2.5 Pro rose by 22.2 points, Gemini 3.1 Pro rose by 9.5 points, and GPT-5.5 rose by 9.3 points, for a total of five models moving upward. The only model moving downward was Claude Sonnet 4.6, which fell by 13 points. The current Top 5 are, in order: Gemini 3.1 Pro at 97.70, Grok 4 at 96.30, GLM-4.6 at 95.00, GPT-o3 at 94.80, and Claude Opus 4.7 at 94.70. The total commitment-adherence score is calculated as an equal-weighted average of the native percentage score on v3 questions and the converted score from v2 anchor questions, with a worst-of-3 sampling standard.
Cause Analysis: Mechanistic Differences Between Pressure Rounds and Constraint Scenarios
WDCD v3 questions use 8–12 rounds of dialogue. They first establish 2–5 hard constraints, then apply social proof, special authorization by authority, salami-slicing, and sunk-cost pressure in sequence, before scoring through a KBV restatement probe and a final-round honest retrospective. S_hold, or commitment-adherence survival, accounts for 60 points, with higher scores awarded the later a breach occurs. Gemini 2.5 Pro and Gemini 3.1 Pro achieved higher S_hold scores under progressive multi-round pressure, indicating that they are better able to maintain constraints in resource-limited and safety-compliance scenarios during rounds 6–10. Claude Sonnet 4.6’s 13-point decline most likely occurred during the R3 pressure phase or the KBV probe stage, suggesting reduced resistance to interference framed as special authorization by authority. In the v2 three-round anchor questions, R3 carries a weight of 2 points; if Sonnet loses points in R3, it directly pulls down the total score. The double-digit gains by Doubao Pro and GPT-5.5 may stem from improved S_recover capability, meaning better recovery after a breach in engineering-specification scenarios.
Implications for Model Selection: Practical Risk Boundaries for Production Workflow Integration
For enterprises integrating AI into production workflows, Gemini 3.1 Pro, with a WDCD score of 97.70, can be used directly in data-boundary and safety-compliance scenarios, with the lowest probability of constraint breach. Grok 4 and GLM-4.6 follow closely and are suitable for resource-limited tasks. Claude Sonnet 4.6’s 13-point decline means that it requires additional guardrails in business-rule scenarios, such as forcibly injecting a KBV restatement check after the fourth round. Although GPT-o3 and Claude Opus 4.7 enter the Top 5, they still trail Gemini 3.1 Pro by more than 3 points, leaving uncertainty in stages where sunk-cost pressure is increased step by step. Enterprises can use models scoring above 95 on WDCD for high-compliance pipelines, while models below 95 should be limited to low-risk internal tools, with human final-round review retained for all models.
Strategic Assessment: Signals of Underestimation and Overestimation
The simultaneous sharp rise of the two Gemini-series models indicates that their constraint memory and honest self-reporting capabilities under v3 multi-round pressure may previously have been underestimated by the market. The isolated decline of Claude Sonnet 4.6 suggests that its commitment-adherence capability may have been overestimated by earlier optimistic expectations, and its S_integrity score should continue to be monitored in the next v3.1 round to see whether it remains at 0. Grok 4 holds steady in second place with 96.30 points but does not appear on this round’s change list, indicating that its commitment-adherence baseline has entered a high plateau; it is worth verifying next round whether it can withstand more aggressive salami-slicing strategies. GLM-4.6 and GPT-o3 sit near the 95-point mark, and any future fluctuation of more than 3 points will directly affect the Top 5 ranking.
The opposite movements of 22.2 points and 13 points between Gemini 3.1 Pro and Claude Sonnet 4.6 have made them the most important comparison pair to track in the next cycle.
Data: YZ Index WDCD Commitment-Adherence Ranking | Run #306 · Change Tracking | Evaluation Methodology
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接