The WDCD v3.1 constraint-adherence test, comparing 11 models across five major constraint scenarios, shows that safety compliance was the lowest-scoring dimension overall, with qwen3-max scoring only 2.19/4, far below the overall median of 3.6/4 in the engineering standards scenario.
Safety Compliance Becomes the Biggest Challenge, with Pressure Rounds Exposing Memory Gaps
In the safety compliance scenario, grok-4 ranked first with 3.86/4, while qwen3-max came last with just 2.19/4, a range of 1.67 points. The v3 item design uses 8–12 consecutive dialogue rounds that first establish constraints and then apply pressure. In the safety compliance items, the “salami slicing” and “special authorization by authority” rounds were the most likely to trigger constraint violations. After pressure was applied in R3, qwen3-max’s score on the KBV restatement probe fell sharply, indicating that its long-term memory of “compliance boundaries” is the first to degrade under multi-round interference. By contrast, the data boundary scenario had a range of only 1.08 points, suggesting that safety compliance constraints are inherently harder for models to internalize.
Resource Limits Show the Highest Differentiation, with glm-4.6 Reaching 3.86/4
In the resource limits scenario, glm-4.6 led with 3.86/4, while gpt-5.5 ranked last with 2.86/4, a range of 1 point. Under the three-round design of the v2 anchor items, R2 interference and R3 pressure accounted for a combined 2 points. glm-4.6 maintained the constraints in both rounds, showing stronger anchoring ability for engineering constraints such as “concurrency quotas” and “timeout fallback.” This scenario had the greatest differentiation, directly reflecting reliability differences among models when handling hard resource constraints in real-world production environments.
Business Rules and Engineering Standards Both Score High, While claude-sonnet-4.6 Shows a 1.07-Point Imbalance
In the business rules scenario, glm-4.6 and grok-4 both scored 4/4, while doubao-pro scored only 2.47/4. The engineering standards scenario was also led by glm-4.6 and gpt-o3, both at 4/4. claude-sonnet-4.6 scored 3.6/4 in engineering standards but only 2.53/4 in business rules, a scenario gap of 1.07 points, meeting the defined threshold for uneven performance. This model adheres relatively steadily to constraints such as “code review standards,” but is more likely to break constraints after sunk-cost pressure in business-rule contexts such as “pricing strategy” and “approval workflows.”
Comparison of Uneven Performance in gemini-2.5-pro and gpt-o3
gemini-2.5-pro scored 3.67/4 in business rules but only 2.51/4 in safety compliance, a gap of 1.16 points; gpt-o3 scored 4/4 in engineering standards but only 2.58/4 in data boundaries, a gap of 1.42 points. Both models showed constraint forgetting during the KBV restatement probe in the v3 items, indicating structural differences in their memory priorities for different types of constraints. Enterprises that involve both data boundaries and engineering standards should add an external validation layer when using gpt-o3.
Specific Recommendations for Enterprise Model Selection
Enterprises integrating AI into production environments should prioritize grok-4 or deepseek-v4-pro for safety compliance scenarios, with scores of 3.86/4 and 3.57/4 respectively; glm-4.6 is the preferred choice for resource limits; and glm-4.6 and grok-4 can both be considered for business rules and engineering standards scenarios. qwen3-max ranked last in both safety compliance and data boundaries, and is not recommended for direct use in scenarios with high compliance requirements. claude-sonnet-4.6 is suitable for engineering-standards-heavy scenarios, but manual review checkpoints should be added in business-rule workflows.
Strategic Assessment
glm-4.6 ranked in the top two across resource limits, business rules, and engineering standards, suggesting that its constraint-adherence capabilities may be underestimated by the market. qwen3-max scored only 2.19/4 in safety compliance, giving it the largest risk exposure; the next evaluation should focus on verifying its constraint recovery mechanism under v3-style multi-round progressive pressure. The overall low scores in the safety compliance scenario indicate that current models still have systemic shortcomings in long-term memory for parallel hard compliance constraints.
Safety compliance is not a moral exam for models; it is a hard metric for whether production systems can avoid long-term failures.
Data source: YZ Index WDCD Constraint-Adherence Leaderboard | Run #311 · Scenario Matrix | Evaluation Methodology
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接