Business Rules Become the Biggest Weakness in WDCD: qwen3-max Scores 1.55 vs deepseek 4

WDCD v3.1 pilot data shows that the business rules scenario is the dimension with the lowest survival rate among the 11 models. Champion deepseek-v4-pro scored 4/4, while last-place qwen3-max managed only 1.55/4, with overall differentiation clearly higher than the other four constraint categories.

Why Business Rules Are a Collective Challenge

The pressure path in the business rules scenario includes three consecutive rounds of anchor interference and salami-slicing rule relaxation. qwen3-max scored only 0.55 in the R3 pressure stage, far below its full score of 4/4 in the engineering specification scenario, a gap of 2.45 points. In contrast, deepseek-v4-pro maintained 4/4 in the same scenario, indicating stronger memory retention for parallel hard constraints. The safety compliance scenario also saw low scores: qwen3-max 1.6/4, doubao-pro 2/4, suggesting that when constraints involve external compliance rather than internal processes, models are more likely to break down under sunk cost pressure.

High Consistency in Data Boundary Scenarios

In the data boundary scenario, seven models — claude-opus-4.7, claude-sonnet-4.6, deepseek-v4-pro, gemini-2.5-pro, glm-4.6, gpt-o3, and grok-4 — all achieved 4/4, with only gpt-5.5 trailing at 3/4. This indicates that current mainstream models have relatively mature compliance capabilities for static hard constraints such as "do not leak training data boundaries," with S_hold scores generally above 55 points.

Intermediate Differentiation in Resource Constraints and Engineering Specifications

In the resource constraints scenario, champion glm-4.6 reached 3.85/4, while last-place qwen3-max scored 2.15/4, a gap of 1.7 points. The engineering specification scenario showed a reversal: claude-sonnet-4.6, gemini-3.1-pro, and qwen3-max tied at 4/4, while doubao-pro scored only 2.15/4. gemini-2.5-pro managed only 2.3/4 in resource constraints but achieved 4/4 in data boundaries, a gap of 1.7 points, revealing significant differences in its processing mechanisms for the two constraint types: "compute resource limits" and "data scope."

Selection Risks of Lopsided Models

claude-sonnet-4.6 scored 4/4 in data boundaries but only 1.8/4 in business rules, a gap of 2.2 points. If a company's core processes involve complex business rule judgments, this model requires additional rule engine safeguards. qwen3-max scored 4/4 in engineering specifications but fell below 1.7/4 in both business rules and safety compliance, indicating a disconnect between its code-level specification adherence and business logic consistency. glm-4.6 scored 4/4 in data boundaries but only 2.3/4 in engineering specifications, making it suitable for data-sensitive scenarios with relatively simple engineering processes.

Specific Recommendations for Production Integration

For enterprises integrating AI into production workflows, claude-opus-4.7 or deepseek-v4-pro can be directly used in data boundary scenarios, with expected low breakage rates. In business rules scenarios, it is recommended to add an independent rule validation layer for models other than deepseek-v4-pro, especially qwen3-max and claude-sonnet-4.6. In safety compliance scenarios, gpt-o3 and grok-4 maintain 4/4 and can be the first choice for compliance-sensitive tasks, but manual review points should still be added after the R3 pressure round.

Low-scoring models in resource constraints scenarios generally have S_recover scores below 6 after continuous pressure, indicating insufficient recovery capability. In production environments, a clear secondary check on token or API call limits should be set.

Strategic Assessment

This data suggests that qwen3-max's engineering specification capability may be overestimated by the market. Its 1.55/4 in business rules stands in sharp contrast to its 4/4 in engineering specifications, warranting more mixed business-engineering constraint questions in the next v3.2 release for verification. deepseek-v4-pro achieved perfect scores in both business rules and data boundaries, suggesting its compliance capability may be underestimated, making it suitable as a multi-scenario baseline model for long-term tracking. The generally low scores in safety compliance scenarios indicate that current models still lack sufficient resistance to "authority override" social engineering pressure. Enterprises should prioritize manual fallback in such scenarios.

WDCD v3.1's worst-of-3 sampling has revealed that relying solely on a model's perfect scores in a single scenario carries significant lopsided risk. The next phase should focus on observing the distribution of S_kbv constraint memory scores in business rules scenarios to determine whether models have truly internalized multiple parallel hard constraints.


Data source: YZ Index WDCD Compliance Leaderboard | Run #233 · Scenario Matrix | Evaluation Methodology

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!