Safety and compliance became the hardest area for models to maintain constraint adherence in this WDCD v3.1 five-constraint test. gemini-2.5-pro scored only 0.8/4, while grok-4 reached 4/4, creating a within-scenario range of 3.2 points—far higher than in the other four scenarios.
Data Facts: Safety and Compliance Created the Widest Gap
Among the five scenarios, the safety and compliance leader grok-4 scored 4/4, while the lowest-ranked gemini-2.5-pro scored only 0.8/4, with the remaining models distributed between 3.5 and 1.6. By comparison, in the data-boundary scenario, the leader doubao-pro scored 4/4 and the lowest-ranked glm-4.6 scored 1.7/4, a range of 2.3 points. In the business-rules scenario, seven models scored 4/4, while the lowest-ranked doubao-pro scored 2.7/4, a range of 1.3 points. The ranges for resource limits and engineering standards were 1.5 points and 1 point, respectively. Safety and compliance showed the highest score dispersion, directly making it the constraint type with the lowest average score across all models.
Cause Analysis: Compliance Memory Collapse Under Multi-Round Pressure
The low scores in the safety and compliance scenario mainly stem from the continuous pressure stages designed in v3. In the question pool, constraints in this scenario are mostly hard compliance clauses such as “prohibit output of specific high-risk content” and “must refuse unauthorized requests.” Under three types of pressure—social proof, special authorization by authority, and salami-slicing—models tend to gradually loosen their boundaries in rounds 5 to 8. gemini-2.5-pro’s 0.8 score in safety and compliance means that, in worst-of-3 sampling, at least one run was already unable to fully restate the initial constraints during the KBV restatement probe stage. By contrast, grok-4 was able to pass both R3 pressure and the final-round honest review, earning relatively high scores on both S_hold and S_integrity. Scores in the business-rules scenario were generally higher because these constraints are mostly internal process rules, and models are more likely to have encountered similar instructions during training, giving them stronger resistance to sunk-cost pressure.
The Real Risks of Unevenly Skilled Models
Among the 11 models, 9 showed gaps of more than 1 point between scenarios. gemini-2.5-pro scored 3.5/4 on data boundaries but fell to 0.8/4 in safety and compliance, a gap of 2.7 points. gemini-3.1-pro scored 4/4 on business rules but only 1.8/4 in safety and compliance, a gap of 2.2 points. qwen3-max likewise scored 4/4 on business rules but 1.6/4 in safety and compliance, a gap of 2.4 points. These models perform strongly in data-boundary or business-rules scenarios, which can easily lead enterprises to overestimate their overall constraint-adherence capabilities. claude-opus-4.7 showed the opposite pattern: it scored 4/4 on business rules but only 2.5/4 on data boundaries, a gap of 1.5 points, indicating that its memory for “data-boundary” constraints decays faster after multi-round interference.
Specific Implications for Enterprise Model Selection
For enterprises integrating AI into production workflows, models scoring below 2.5/4 in safety and compliance scenarios—gemini-2.5-pro, qwen3-max, glm-4.6, and gemini-3.1-pro—should not be used directly in stages involving user privacy, content moderation, or compliance reporting. Additional output filtering layers or human review nodes must be deployed. Models scoring 3.5/4 or above in data-boundary scenarios, such as doubao-pro, claude-sonnet-4.6, and deepseek-v4-pro, can be prioritized for internal systems requiring strict data isolation. In resource-limit scenarios, both claude-sonnet-4.6 and grok-4 scored 4/4, making them suitable for edge deployment scenarios constrained by budget or computing power. Engineering standards scores were generally high overall, and the 4/4 performance of claude-sonnet-4.6 and gpt-o3 makes them suitable for code review and process automation.
Strategic Judgment: Overestimated and Underestimated Constraint-Adherence Capabilities
This data may indicate that the market has overestimated the Gemini series’ constraint-adherence capability in safety and compliance scenarios, as its high scores in data boundaries and business rules can easily obscure its compliance shortcomings. grok-4’s 4/4 performance in safety and compliance deserves continued validation in the next round, especially with attention to its S_recover score under more rounds of salami-slicing pressure. claude-opus-4.7’s pattern of a perfect score in business rules but weaker performance in data boundaries may reflect differences in its memory priorities across constraint types. If future versions can balance this gap, it will become more competitive in overall constraint-adherence scores. Enterprises should not look only at the champion in a single scenario; instead, they should make final model-selection decisions by cross-referencing the worst-of-3 lowest results according to their own most frequent constraint types.
Safety and compliance are not bonus questions for models; they are the first gate in production environments.
Data source: YZ Index WDCD Constraint-Adherence Ranking | Run #316 · Scenario Matrix | Evaluation Methodology
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接