GPT-o3 Stages Comeback with 9.5-Point Rise, GLM-4.6 Plunges 14.9 — Five Models Reshuffled on WDCD Compliance Leaderboard

This round of WDCD v3.1 testing shows GPT-o3 up 9.5 points from Run #253, Gemini 2.5 Pro up 7.6 points, while GLM-4.6 is down 14.9 points, Claude Sonnet 4.6 is down 10.8 points, and Claude Opus 4.7 is down 5.9 points. Among the 11 models evaluated, 2 rose and 3 fell.

Data Facts: Specific Scores and Ranking Shifts

In the current Top 5, Grok 4 remains first with WDCD=97.50, GPT-o3 ranks second with WDCD=95.20, DeepSeek V4 Pro ranks third with WDCD=91.10, Claude Opus 4.7 ranks fourth with WDCD=86.70, and Gemini 3.1 Pro ranks fifth with WDCD=84.50. With this round's rise, GPT-o3 has moved into the top two, while GLM-4.6, due to its 14.9-point drop, has shifted noticeably backward from its previous position. The sampling protocol is worst-of-3: each question is run three times and the worst result is taken; v3 questions and v2 anchor questions are equally weighted.

Cause Analysis: Possible Mechanisms in Pressure Rounds and Constraint Scenarios

WDCD v3 questions use 8–12 rounds of dialogue, first establishing 2–5 parallel hard constraints, then sequentially applying social proof, authoritative approval, salami-slicing, and sunk-cost pressure, followed by a KBV paraphrase probe and a final-round honesty review. GLM-4.6's 14.9-point drop is the most significant, possibly stemming from lower S_hold compliance survival scores during the mid-to-late rounds of "salami-slicing" and "sunk-cost" escalation. Claude Sonnet 4.6's 10.8-point drop and Opus 4.7's 5.9-point drop suggest more deductions in the S_kbv constraint memory or S_recover breach-recovery stage in multi-constraint parallel scenarios. GPT-o3's 9.5-point rise and Gemini 2.5 Pro's 7.6-point rise may reflect more stable performance in the final-round S_integrity honest self-report and the R3 pressure stage of the v2 three-round anchor questions.

The five constraint scenarios include data boundaries, resource limits, business rules, security compliance, and engineering specifications. For declining models, the worst-of-3 worst-case performance in the security compliance and business rules scenarios dragged down total scores, while rising models scored relatively higher on S_hold in the resource limits and engineering specifications scenarios. The v2 anchor questions have a maximum of 4 points (R1:1 + R2:1 + R3:2), and this round's changes are also reflected in the R3 pressure stage.

Selection Implications: Actual Risk Boundaries for Production Integration

For enterprises integrating AI into production workflows, Grok 4 WDCD=97.50 and GPT-o3 WDCD=95.20 mean that in scenarios requiring strict execution of multiple parallel hard constraints, these models can be considered for direct use with a relatively low likelihood of breaking compliance. Claude Opus 4.7 WDCD=86.70 remains in the top five, but its 5.9-point drop from Run #253 suggests that under sustained multi-round social proof and authoritative approval pressure, additional manual review or secondary confirmation mechanisms are needed. After the 14.9-point decline, GLM-4.6's applicability in security compliance and data boundary scenarios needs to be reassessed; external guardrails or rule-engine interception are recommended for these scenarios.

Specific assessment: in resource limits and engineering specifications scenarios, the upward data for GPT-o3 and Gemini 2.5 Pro supports a higher degree of trust; in business rules and security compliance scenarios, the downward data for the two Claude models and GLM-4.6 requires enterprises to implement stricter output validation.

Strategic Assessment: Market Valuation Deviation of Compliance Capability

Judging from this round's data, Grok 4's leading position at WDCD=97.50 and GPT-o3's rise to WDCD=95.20 indicate that their compliance capabilities may have been previously underestimated by the market. GLM-4.6's 14.9-point drop and Claude Sonnet 4.6's 10.8-point drop show a gap between their actual performance under mid-to-late round pressure in v3 questions and some market expectations. DeepSeek V4 Pro holds third place at WDCD=91.10, and Gemini 3.1 Pro ranks fifth at WDCD=84.50; no near-term adjustment to their baseline assessments is needed.

Signals worth verifying in the next round: the worst-of-3 stability of GPT-o3 and Gemini 2.5 Pro in the KBV paraphrase probe and S_integrity honest self-report stages, as well as GLM-4.6's recovery capability in the R3 pressure stage. The analysis is based solely on Run #253, the only comparable dataset for this round, with no additional assumptions introduced.

The compliance score is not a static label; it is the real moat of an enterprise under genuine multi-round pressure.

Data source: YZ Index WDCD Compliance Leaderboard | Run #263 · Change Tracking | Evaluation Methodology

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!