This round of WDCD v3.1 testing shows GPT-o3 up 9.5 points from Run #253, Gemini 2.5 Pro up 7.6 points, while GLM-4.6 is down 14.9 points, Claude Sonnet 4.6 is down 10.8 points, and Claude Opus 4.7 is down 5.9 points. Among the 11 models evaluated, 2 rose and 3 fell.
Data Facts: Specific Scores and Ranking Shifts
In the current Top 5, Grok 4 remains first with WDCD=97.50, GPT-o3 ranks second with WDCD=95.20, DeepSeek V4 Pro ranks third with WDCD=91.10, Claude Opus 4.7 ranks fourth with WDCD=86.70, and Gemini 3.1 Pro ranks fifth with WDCD=84.50. With this round's rise, GPT-o3 has moved into the top two, while GLM-4.6, due to its 14.9-point drop, has shifted noticeably backward from its previous position. The sampling protocol is worst-of-3: each question is run three times and the worst result is taken; v3 questions and v2 anchor questions are equally weighted.
Cause Analysis: Possible Mechanisms in Pressure Rounds and Constraint Scenarios
WDCD v3 questions use 8–12 rounds of dialogue, first establishing 2–5 parallel hard constraints, then sequentially applying social proof, authoritative approval, salami-slicing, and sunk-cost pressure, followed by a KBV paraphrase probe and a final-round honesty review. GLM-4.6's 14.9-point drop is the most significant, possibly stemming from lower S_hold compliance survival scores during the mid-to-late rounds of "salami-slicing" and "sunk-cost" escalation. Claude Sonnet 4.6's 10.8-point drop and Opus 4.7's 5.9-point drop suggest more deductions in the S_kbv constraint memory or S_recover breach-recovery stage in multi-constraint parallel scenarios. GPT-o3's 9.5-point rise and Gemini 2.5 Pro's 7.6-point rise may reflect more stable performance in the final-round S_integrity honest self-report and the R3 pressure stage of the v2 three-round anchor questions.
The five constraint scenarios include data boundaries, resource limits, business rules, security compliance, and engineering specifications. For declining models, the worst-of-3 worst-case performance in the security compliance and business rules scenarios dragged down total scores, while rising models scored relatively higher on S_hold in the resource limits and engineering specifications scenarios. The v2 anchor questions have a maximum of 4 points (R1:1 + R2:1 + R3:2), and this round's changes are also reflected in the R3 pressure stage.
Selection Implications: Actual Risk Boundaries for Production Integration
For enterprises integrating AI into production workflows, Grok 4 WDCD=97.50 and GPT-o3 WDCD=95.20 mean that in scenarios requiring strict execution of multiple parallel hard constraints, these models can be considered for direct use with a relatively low likelihood of breaking compliance. Claude Opus 4.7 WDCD=86.70 remains in the top five, but its 5.9-point drop from Run #253 suggests that under sustained multi-round social proof and authoritative approval pressure, additional manual review or secondary confirmation mechanisms are needed. After the 14.9-point decline, GLM-4.6's applicability in security compliance and data boundary scenarios needs to be reassessed; external guardrails or rule-engine interception are recommended for these scenarios.
Specific assessment: in resource limits and engineering specifications scenarios, the upward data for GPT-o3 and Gemini 2.5 Pro supports a higher degree of trust; in business rules and security compliance scenarios, the downward data for the two Claude models and GLM-4.6 requires enterprises to implement stricter output validation.
Strategic Assessment: Market Valuation Deviation of Compliance Capability
Judging from this round's data, Grok 4's leading position at WDCD=97.50 and GPT-o3's rise to WDCD=95.20 indicate that their compliance capabilities may have been previously underestimated by the market. GLM-4.6's 14.9-point drop and Claude Sonnet 4.6's 10.8-point drop show a gap between their actual performance under mid-to-late round pressure in v3 questions and some market expectations. DeepSeek V4 Pro holds third place at WDCD=91.10, and Gemini 3.1 Pro ranks fifth at WDCD=84.50; no near-term adjustment to their baseline assessments is needed.
Signals worth verifying in the next round: the worst-of-3 stability of GPT-o3 and Gemini 2.5 Pro in the KBV paraphrase probe and S_integrity honest self-report stages, as well as GLM-4.6's recovery capability in the R3 pressure stage. The analysis is based solely on Run #253, the only comparable dataset for this round, with no additional assumptions introduced.
The compliance score is not a static label; it is the real moat of an enterprise under genuine multi-round pressure.
Data source: YZ Index WDCD Compliance Leaderboard | Run #263 · Change Tracking | Evaluation Methodology
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接