WDCD Cross-Scenario Review: Safety Compliance Falls to 1.1, 15 Models Show Imbalance Gaps of Up to 2.9 Points

WDCD v3.1's cross-evaluation of five scenarios shows that the safety compliance scenario is the low point in adherence capability for the 15 models, with qwen3-max scoring only 1.1/4, far below the lowest score of 3/4 in the engineering standards scenario.

Safety Compliance Collapses Across the Board, Data Boundaries Close Behind

In the safety compliance scenario, co-leaders glm-4.6 and grok-4 both scored 4/4, but bottom-ranked qwen3-max managed only 1.1/4, and gpt-6-luna just 1.2/4. The spread between the highest and lowest scores within the same scenario reached 2.9 points, the widest among the five scenarios. The data boundaries scenario was similarly severe: qwen3-max again came last at 1.5/4, while claude-opus-4.7 and doubao-pro tied at 4/4. By contrast, the lowest score in the engineering standards scenario was still 3/4, indicating stronger overall resilience.

Pressure Rounds and Constraint Type Determine Score Differences

The low scores in the safety compliance scenario stem mainly from two types of continuous pressure in the v3 question set: "authority-granted special approval" and "salami slicing." qwen3-max broke its commitments repeatedly as early as the R2 interference round, and after R3 pressure its S_hold adherence survival score fell sharply. The data boundaries scenario exposed memory decay in the KBV paraphrasing probe: gpt-6-astra dropped directly from 4/4 to 2.35/4, indicating insufficient ability to sustain the "data boundary" hard constraint over time. Most models scored 4/4 in the business rules scenario, suggesting this type of constraint aligns more closely with pretraining and that models can still maintain S_recover recovery capability after pressure.

Ten Models Show Clear Imbalance, with Gaps Exceeding 1 Point

claude-sonnet-4.6 scored 4/4 on business rules but only 2.4/4 on safety compliance, a scenario gap of 1.6 points. gemini-2.5-pro scored 4/4 on business rules and only 2.15/4 on safety compliance, a gap of 1.85 points. gpt-6-luna scored 4/4 on business rules and 1.2/4 on safety compliance, a gap of 2.8 points. qwen3-max scored 4/4 on engineering standards and 1.1/4 on safety compliance, a gap of 2.9 points. These imbalances show that models' adherence mechanisms differ structurally across constraint types, rather than reflecting a simple difference in overall capability.

Scenario-Based Selection Advice for Enterprises Integrating AI into Production Workflows

Enterprises bringing AI into safety compliance workflows should prioritize glm-4.6 or grok-4, both of which score 4/4 in that scenario and show stable S_integrity honest self-report scores. For the data boundaries scenario, claude-opus-4.7 and doubao-pro are recommended, while qwen3-max should be avoided. In the resource limits scenario, doubao-pro, gemini-3.1-pro, and glm-4.6 all score 4/4 and can be used directly. In the business rules scenario, nearly all models reach 4/4, but manual review should still be added during the R3 pressure round. The engineering standards scenario is the most stable overall; gemini-2.5-pro, though lowest at 3/4, remains acceptable.

Strategic Judgment: Some Models' Adherence Capability Is Overestimated by the Market

qwen3-max scored 4/4 on engineering standards yet came last in both safety compliance and data boundaries; its adherence capability may be overestimated by the market. claude-opus-4.7 scored 4/4 in all three of the data boundaries, business rules, and engineering standards scenarios, giving it the most balanced overall adherence survival capability, but it scored only 3.25/4 on resource limits and needs supplementary external constraints in that scenario. gpt-6-luna's extreme low of 1.2/4 on safety compliance signals high risk in compliance-critical production environments. The next round of testing could focus on the S_kbv constraint memory score in the safety compliance scenario to verify whether low-scoring models suffer cascading commitment breaches due to KBV probe failure.

The 1.1 score in the safety compliance scenario is no accident — it exposes the true adherence boundaries of models under multi-round pressure in the v3 question set.

When selecting models, enterprises must not rely on a single-scenario champion; they must match models to actual constraint types and deploy independent guardrails for the safety compliance and data boundaries scenarios. The WDCD v3.1 data clearly shows that adherence capability is no longer a general property of models, but a scenario-specific vulnerability.


Data sources: YZ Index WDCD Adherence Leaderboard | Run #365 · Scenario Matrix | Evaluation Methodology

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!