Claude Sonnet 4.6 Surges 15 Points, GLM-4.6 Plunges 15.3: WDCD Compliance Polarization

Claude Sonnet 4.6 rose 15 points in the latest WDCD v3.1 test compared to Run #227, while GLM-4.6 dropped 15.3 points in the same period. This symmetric positive-negative shift emerged as the most prominent compliance fluctuation signal among the 11 evaluated models.

Data Facts: Precise Distribution of 3 Rises and 4 Falls

This round used a worst-of-3 measurement, covering 17 progressive pressure questions from the v3.1 question pool and 8 v2 anchor questions. Only 3 models rose, while 4 declined. The specific values are: Claude Opus 4.7 up 5.1 points, Claude Sonnet 4.6 up 15 points, Gemini 3.1 Pro up 6.5 points; DeepSeek V4 Pro down 9.7 points, Doubao Pro down 6.8 points, GLM-4.6 down 15.3 points, and Grok 4 down 7.1 points. In the current Top 5, GPT-o3 remains first with 94.00 points, Grok 4 is second with 87.90 points, Claude Opus 4.7 third with 87.60 points, Gemini 3.1 Pro fourth with 87.30 points, and DeepSeek V4 Pro fifth with 84.30 points.

Cause Analysis: Mechanistic Differences in Pressure Rounds and Constraint Scenarios

WDCD v3.1 design includes 8-12 consecutive rounds of social approval, authority special approval, salami slicing, and escalating sunk costs after the commitment, as well as KBV repetition probes and a final honesty self-report. Claude Sonnet 4.6's 15-point gain most likely comes from its late-breakthrough performance in the S_hold compliance survival 60-point metric, i.e., maintaining constraint boundaries under multiple rounds of salami-slicing pressure. In contrast, GLM-4.6's 15.3-point decline is more likely concentrated in the S_recover breach recovery 10-point and S_integrity honest self-report 15-point metrics, suggesting it is more prone to memory loss of constraints that cannot be recovered during the R3 pressure stage of v2 three-round anchor questions.

Among the five types of constraint scenarios, safety compliance and engineering specification scenarios have the steepest pressure gradient. Claude Sonnet 4.6's score improvement may reflect its enhanced resistance to authority special approval-type interference, while GLM-4.6's decline may stem from sunk cost pressure failure in resource-constrained scenarios. Although Grok 4 still ranks second, its 7.1-point drop, together with DeepSeek V4 Pro's 9.7-point drop, indicates fragility in the S_kbv constraint memory 15-point metric in data boundary scenarios.

Selection Implications: Practical Risk Boundaries for Production Pipeline Integration

For enterprises integrating AI into production pipelines, GPT-o3's 94.00 points mean it can be used directly in safety compliance scenarios; its high S_hold score indicates a low probability of breakthrough under multiple rounds of business rule pressure. Grok 4 and Claude Opus 4.7 are tied in the 87-point range, suitable for resource-constrained scenarios but require additional secondary confirmation guardrails in engineering specifications. The gap between Gemini 3.1 Pro's 87.30 points and DeepSeek V4 Pro's 84.30 points shows that the latter's compliance stability in data boundary scenarios is already 6 points lower than the former. Enterprises planning to call DeepSeek V4 Pro to process sensitive data should add manual review nodes at the KBV repetition probe positions corresponding to v3 questions.

Claude Sonnet 4.6's 15-point increase suggests it can be used for initial trials in business rule scenarios, but GLM-4.6's 15.3-point decline is a clear warning: any production pipeline relying on its long-term constraint memory must immediately install an external state machine to avoid cumulative violations caused by escalating sunk costs.

Strategic Judgment: Underestimated and Overestimated Compliance Signals

The 15-point gain of Claude Sonnet 4.6 compared to Claude Opus 4.7's 5.1-point gain shows an internal divergence, indicating that the Sonnet version may have received targeted optimizations for the continuous pressure rounds of v3.1. This signal warrants focused verification of its S_recover performance in the next v3.1 question pool. GLM-4.6's 15.3-point decline, exceeding Grok 4's 7.1-point decline, shows that its compliance capability has been relatively overestimated by the market. Enterprises should lower their trust weight for GLM-4.6 in engineering specification scenarios during model selection.

GPT-o3 leads the second place by 6.1 points with 94.00 points, the most stable lead in this data. Although DeepSeek V4 Pro still ranks fifth, its 9.7-point drop has widened the gap with Gemini 3.1 Pro to 3 points. If it continues to lose points in safety compliance scenarios in the future, it may face the risk of dropping out of the Top 5.

Compliance capability divergence is not a simple function of model size, but a true response to multi-round pressure escalation mechanisms. The data in this round has clearly marked who is regressing.

Data Source: YZ Index WDCD Compliance Ranking | Run #233 · Change Tracking | Evaluation Methodology

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!