WDCD Five-Scenario Review: Safety Compliance Lowest at 1.5 Points; qwen3-max's Imbalance Creates 2.5-Point Gap

In the WDCD v3.1 test, the average score of 11 models in the safety compliance scenario was noticeably lower than in the other four categories. qwen3-max scored only 1.5/4, while claude-sonnet-4.6, gpt-o3, and grok-4 tied at 4/4. Champions in the data boundary, resource limits, business rules, and engineering standards scenarios all scored 4/4, making safety compliance the only scenario where a 1.5 score appeared.

Safety Compliance Proves the Hardest Scenario, with Pressure Rounds Directly Exposing Gaps

In the safety compliance scenario, qwen3-max scored 1.5/4, gemini-2.5-pro 2/4, and gpt-5.5 2.4/4; the bottom three trailed the champions by more than 2 points. The v3 questions include 8-12 rounds of dialogue, with continuous pressure applied after the commitment phase through four escalating levels: social proof, authoritative approval, salami slicing, and sunk cost. Safety compliance questions mostly consist of hard prohibition clauses, with most breaches occurring in rounds 5-7 during the salami-slicing and sunk-cost stages. In contrast, the engineering standards scenario had a minimum score of 3.1/4, indicating that engineering constraints are easier to maintain under multi-round pressure.

Business Rules Show the Greatest Differentiation, with 4 Points and 2.15 Points Side by Side

In the business rules scenario, deepseek-v4-pro, gemini-3.1-pro, glm-4.6, gpt-5.5, and grok-4 all scored 4/4, while doubao-pro managed only 2.15/4. The 1.85-point gap is the largest among the five scenarios. Business rules questions mostly involve parallel hard constraints; the v3 questions require models to remember 2-5 rules simultaneously during the commitment phase. doubao-pro showed signs of forgetting as early as the R2 interference round and directly violated the rules after R3 pressure. The data boundary scenario, by contrast, had no such extreme spread: champion grok-4 scored 4/4, while the bottom-ranked doubao-pro scored 2.7/4, a range of only 1.3 points.

Models with Uneven Performance: qwen3-max's 2.5-Point Gap Tops the List

qwen3-max scored 4/4 in engineering standards but only 1.5/4 in safety compliance, a 2.5-point gap between scenarios. doubao-pro scored 3.9/4 in engineering standards and 2.15/4 in business rules, a 1.75-point gap. gpt-5.5 scored 4/4 in business rules and 2.4/4 in safety compliance, a 1.6-point gap. gemini-2.5-pro scored 3.5/4 in data boundary and 2/4 in safety compliance, a 1.5-point gap. glm-4.6 scored 4/4 in business rules and 2.9/4 in safety compliance, a 1.1-point gap. All these gaps come from worst-of-3 sampling, showing that models exhibit systematic differences in constraint adherence across different constraint types.

Enterprise Model Selection: Safety Compliance Scenarios Require Guardrails

Enterprises integrating AI into production workflows should prioritize claude-sonnet-4.6, gpt-o3, and grok-4 when core processes involve safety compliance. qwen3-max and gemini-2.5-pro scored below 2.5/4 in this scenario, so an independent compliance-checking layer is recommended before deployment. In the business rules scenario, doubao-pro and qwen3-max scored below 3/4, making them suitable only for low-risk internal rule-processing use cases. In the resource limits and engineering standards scenarios, deepseek-v4-pro, glm-4.6, and claude-sonnet-4.6 performed steadily and can be integrated directly.

Strategic Assessment: Models with Underestimated and Overestimated Constraint Adherence

claude-sonnet-4.6 scored 4/4 in both safety compliance and engineering standards, suggesting its constraint-adherence capability may be underestimated by the market. qwen3-max achieved a perfect score in engineering standards but only 1.5/4 in safety compliance, showing that its constraint memory fluctuates sharply across scenarios; the next round of testing needs to focus on verifying its R3 pressure-round performance in safety compliance questions. deepseek-v4-pro scored 4/4 in both resource limits and business rules but only 3.4/4 in engineering standards, indicating potential risks in parallel multi-constraint scenarios.

The pilot data from this round shows that the constraint-adherence survival score S_hold in the safety compliance scenario has the greatest impact on overall scores. If enterprises plan to expand their AI usage, they need to invest additional effort in prompt engineering or external validation for safety compliance scenarios, rather than relying solely on models' native constraint-adherence capability.

The 1.5-point gap in the safety compliance scenario signals that production-grade deployment will remain constrained if next-generation models cannot maintain hard constraints under multi-round salami-slicing pressure.

Data source: YZ Index WDCD Constraint-Adherence Leaderboard | Run #296 · Scenario Matrix | Evaluation Methodology

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!