WDCD Comparative Review: Safety Compliance Lowest at 1.8 Points, Engineering Standards Full 4 Across the Board

In the WDCD v3.1 compliance test, the safety compliance scenario posted the lowest collective scores: gpt-5.5 and qwen3-max both managed only 1.8/4, while the lowest score among 11 models in the engineering standards scenario reached 3.2/4 — the gap is unmistakable.

Why Safety Compliance Is the Hardest Scenario

In the safety compliance scenario, deepseek-v4-pro, gpt-o3, and grok-4 each took 4/4, while the remaining models fell off steeply. gpt-5.5 and qwen3-max bottomed out at 1.8/4, and claude-sonnet-4.6 managed only 3.2/4. The pressure rounds in this scenario include continuous salami-slicing of compliance boundaries and authority special approvals. Models are most prone to breaking down after the third round, causing substantial losses in the S_hold compliance survival score. By contrast, the engineering standards scenario imposes mostly static code style and version-locking constraints with much lower pressure intensity, which is why all models clustered within the 3.2-4/4 range.

Scenarios with the Greatest and Least Differentiation

The safety compliance scenario shows the greatest differentiation, with scores plunging from 4/4 to 1.8/4 — a spread of 2.2 points. The business rules scenario also widens the gap, with deepseek-v4-pro and gemini-3.1-pro taking 4/4 while doubao-pro managed only 1.7/4. The engineering standards scenario shows the least differentiation, with all 11 models clustered between 3.2 and 4/4, a maximum spread of just 0.8 points. This indicates that engineering standards constraints are nearing saturation for current models, while safety compliance remains in a clear capability gap.

Imbalances Across Scenarios and Their Causes

claude-opus-4.7 took 4/4 in both engineering standards and resource constraints but only 2.5/4 in data boundaries, a 1.5-point gap. Its strengths are concentrated in resource quotas and code standards, while its weakness appears in long-term memory under hard data boundary constraints. gpt-5.5 scored 4/4 in engineering standards but only 1.8/4 in safety compliance, a 2.2-point gap, indicating insufficient recovery under multi-round compliance pressure. gemini-3.1-pro scored 4/4 in business rules but only 2.4/4 in safety compliance, a 1.6-point gap, suggesting faster memory decay in safety-related KBV recall probe rounds. doubao-pro scored 3.4/4 in engineering standards but 1.7/4 in both business rules and data boundaries, a 1.7-point gap, with an overall low compliance baseline.

Scenario-Based Recommendations for Enterprise Production Integration

For enterprises integrating AI into production workflows, gpt-o3 (4/4) or claude-sonnet-4.6 (3.8/4) should be prioritized in data boundary scenarios, with secondary validation added at the interface layer. In resource constraint scenarios, claude-opus-4.7 or grok-4 (both 4/4) can be deployed directly with minimal guardrail needs. For business rules scenarios, deepseek-v4-pro and gemini-3.1-pro (both 4/4) are recommended, though doubao-pro should be reinforced with a rules engine as a fallback. The safety compliance scenario carries the highest risk: deepseek-v4-pro and gpt-o3 serve as the primary choices, while all other models must include human review checkpoints. In engineering standards scenarios, nearly every model is usable, and additional guardrails yield the smallest marginal benefit.

Strategic Assessment: Signals of Underestimated and Overestimated Capabilities

deepseek-v4-pro scored 4/4 across safety compliance, business rules, and engineering standards, with only resource constraints at 2.7/4, suggesting its overall compliance capability may be underestimated by the market. claude-opus-4.7 earned full marks in resource constraints and engineering standards but only 2.5/4 in data boundaries, revealing a shortfall in memory stability under parallel hard constraints. gpt-5.5 scored full marks in engineering standards but 1.8/4 in safety compliance, indicating that its S_recover capability under high-intensity compliance pressure has not yet reached the level implied by its marketing. The next verification cycle should focus on rounds 8-10 of the KBV recall probes in the safety compliance scenario, to observe whether the 1.8-point models can lift their S_integrity scores through additional fine-tuning.

The 1.8 score in the safety compliance scenario is no accident — it is the inevitable result of memory fragmentation after multi-round salami-slicing pressure.

When selecting models, enterprises must not infer reliability across all scenarios from a full score in any single one; they must match each item against the WDCD five-dimensional matrix. Engineering standards are approaching their ceiling, while safety compliance remains the critical bottleneck determining production readiness.


Data source: YZ Index WDCD Compliance Leaderboard | Run #263 · Scenario Matrix | Evaluation Methodology

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!