In this round of the WDCD v3.1 pilot, GLM-4.6 rose 26 points from Run #346, GPT-5.5 rose 10.5 points, and the other 13 evaluated models recorded no declines. GLM-4.6 currently has a WDCD score of 87.55, jumping to third place; Grok 4 remains first at 95.69, while Gemini 3.1 Pro is second at 88.14.
Data Fact: Only Two Models Moved Positively
Among the 15 models, only GLM-4.6 and GPT-5.5 showed positive change. GLM-4.6 rose 26 points from its previous level to reach 87.55 and entered the Top 5; GPT-5.5 gained 10.5 points but did not enter the top five. The v3 question pool contains 17 multi-turn progressive pressure questions and 8 v2 anchor questions, with sampling using a worst-of-3 standard. The total compliance score is derived from the equal-weighted average of the v3 percentage score and the converted value of the v2 anchor questions.
Cause Analysis: S_hold Differences Under Multi-Turn Pressure
GLM-4.6's 26-point increase most likely comes from the 8–12 turn dialogue structure of the v3 questions. The v3 questions first set 2–5 parallel hard constraints, then sequentially apply social proof, authority override, salami slicing, and sunk cost pressure, and finally conduct a KBV restatement probe and a final-round honest debrief. S_hold has a weight of 60 points, and the later a constraint is broken, the higher the score. GLM-4.6 may have held out longer during the continuous pressure phase in rounds 6–9, leading to a significant increase in its S_hold score. GPT-5.5's 10.5-point increase is smaller, possibly reflecting improved performance only in certain constrained scenarios (such as resource limits or engineering standards) in the R3 round, while the v2 anchor questions' three-round scores (R1:1+R2:1+R3:2) contribute relatively little.
There is still an 8.14-point gap between GLM-4.6's WDCD 87.55 and Grok 4's WDCD 95.69, indicating that its constraint memory and recovery from constraint breaches under the worst sample still lag behind.
Selection Implications: Scenario Boundaries for Integration into Production Workflows
For enterprises integrating AI into production workflows, GLM-4.6's improved compliance score means they can reduce some guardrail strength in data boundary and security compliance scenarios. Enterprises can use GLM-4.6 directly in low-risk workflows such as internal knowledge base retrieval and compliance document generation without adding manual review at every turn. However, in scenarios where resource limits and business rules run in strict parallel, final human review should still be retained, because worst-of-3 sampling shows that its worst run can still break constraints late. Grok 4's leading position at WDCD 95.69 makes it suitable as the preferred model for high-value transactions or permission allocation, while Claude Opus 4.7 at WDCD 86.34 is suitable as a secondary review node.
- High-compliance scenarios: prioritize GLM-4.6 or Grok 4 to reduce secondary prompt engineering costs
- Multi-model parallel pipelines: use Gemini 3.1 Pro WDCD 88.14 for intermediate transformation steps
- DeepSeek V4 Pro WDCD 86.62 is suitable for low-sensitivity data processing
Strategic Judgment: A Signal That Compliance Ability Is Underestimated
GLM-4.6's single-round 26-point jump suggests that its S_integrity and S_recover abilities under v3 multi-turn progressive pressure may have been underestimated by the market. DeepSeek V4 Pro and Claude Opus 4.7 scored 86.62 and 86.34, respectively, a gap of only 0.28 points, indicating that their actual performance under current constraint scenarios is close. The next round should focus on determining whether GLM-4.6's S_kbv score in the KBV restatement probe remains stable, and whether GPT-5.5 can replicate its 10.5-point increase on more v3 questions. Grok 4's high score of 95.69 still constitutes a clear moat, but whether its lead will narrow due to single-round breakthroughs by other models is worth continued tracking.
GLM-4.6's 26-point increase and GPT-5.5's 10.5-point increase together point to one fact: at the current stage, improvements in compliance ability mainly occur in the middle rounds of multi-turn pressure, rather than in single-turn anchor questions. When selecting models, enterprises should treat WDCD scores as a dynamic threshold rather than a static ranking.
Data source: YZ Index WDCD Compliance Leaderboard | Run #360 · Change Tracking | Evaluation Methodology
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接