Business Rules Score Lowest at 1.83 Points, Grok-4 Safety Compliance at 3.86: Who Is Least Reliable in Five-Scenario Compliance Testing?

The WDCD v3.1 five-constraint scenario comparative evaluation shows that business rules scenarios scored the lowest across all models, with Doubao-Pro at only 1.83/4, far below Grok-4's perfect 4/4. Safety compliance scenarios also recorded an extremely low score of 1.87/4, with these two scenarios jointly forming the hardest pressure zone to pass in current compliance testing.

Business Rules and Safety Compliance Tie as the Hardest Scenarios

In the score distribution across the five scenarios, the business rules scenario's bottom score of 1.83/4 and the safety compliance scenario's bottom score of 1.87/4 are both significantly lower than the data boundary scenario's 2.8/4 and the engineering standards scenario's 3.46/4. The business rules scenario requires models to maintain compliance under 2-5 parallel hard constraints while facing four escalating rounds of pressure—social conformity, authoritative special approval, salami slicing, and sunk cost. Doubao-Pro broke down multiple times as early as the R3 pressure round, causing substantial loss in its S_hold compliance survival score. The safety compliance scenario similarly uses 8-12 rounds of dialogue, and even after the KBV paraphrase probe, some models still scored 0 in the final round's honest self-reporting stage, indicating deficiencies in both constraint memory and breakdown recovery capabilities.

Business Rules Is the Scenario with the Greatest Differentiation

The business rules scenario's gap between the highest score of 4/4 and the lowest score of 1.83/4 reaches 2.17 points, far exceeding the engineering standards scenario's 0.54-point gap between 4/4 and 3.46/4. Grok-4 achieved the only perfect score in this scenario, with Gemini-2.5-Pro, GLM-4.6, and GPT-5.5 tied at 3.67/4, and Qwen3-Max at only 2.87/4. The gap primarily stems from the R2 interference and R3 pressure rounds—Grok-4 maintained multiple parallel constraints under continuous pressure, while Doubao-Pro abandoned its initial commitments during the sunk cost escalation stage, with an S_recover breakdown recovery score near zero.

The Real Risks of Lopsided Models

Doubao-Pro scored 3.47/4 in the engineering standards scenario but only 1.83/4 in the business rules scenario, a 1.64-point gap between scenarios. Qwen3-Max scored 3.71/4 in engineering standards but only 2.1/4 in safety compliance, a 1.61-point gap. These models perform stably on engineering standards constraints (such as code formatting and version control) but are prone to breaking commitments in business rules or safety compliance scenarios involving multiple stakeholders. For enterprises integrating these models into production workflows, Doubao-Pro is suitable for internal code review tools, but if used for customer contract terms or compliance approval processes, additional secondary manual review or a rules engine fallback is required.

Model Strengths and Weaknesses by Scenario and Selection Implications

Grok-4 ranked first or second in all three scenarios—resource constraints at 3.64/4, safety compliance at 3.86/4, and business rules at 4/4—demonstrating high S_hold and S_integrity scores under multi-round progressive pressure, making it suitable for financial risk control or supply chain systems that must satisfy both resource caps and compliance boundaries. Claude-Sonnet-4.6 scored 4/4 in engineering standards and Claude-Opus-4.7 scored 3.6/4 in data boundaries. Both Claude models excel in engineering standards and data boundary scenarios, but Claude-Opus-4.7 scored only 3.01/4 in the resource constraints scenario, lower than Grok-4's 3.64/4, indicating weaker recovery capability under resource quota constraints.

DeepSeek-V4-Pro scored 3.6/4 in data boundaries and 4/4 in engineering standards, making it suitable for research platforms with strict requirements on data boundaries and code standards. However, its safety compliance score of only 3.39/4 falls below Grok-4's 3.86/4, so deployments need an additional secondary confirmation mechanism after the KBV paraphrase probe in safety compliance scenarios.

Strategic Assessment and Next-Round Validation Signals

Current data indicates that Grok-4's compliance capability is underestimated under multi-scenario pressure, while Doubao-Pro and Qwen3-Max's engineering standards scores may mask their systemic weaknesses in business rules and safety compliance scenarios. The engineering standards scenario has the highest overall scores and the lowest differentiation, suggesting the current v3 question pool does not apply sufficient pressure to this scenario. The next round could add combined pressure rounds of sunk cost and authoritative special approval to improve differentiation. When selecting models, for production workflows involving customer contracts, compliance approval, or resource quota allocation, prioritize Grok-4 while retaining manual final review. For scenarios used only for internal engineering standards checks, Claude-Sonnet-4.6 or DeepSeek-V4-Pro may be considered to reduce costs.

The low scores in business rules and safety compliance are not incidental model errors, but the inevitable result of S_hold and S_kbv collapsing simultaneously under multi-round parallel constraints.

Pilot-phase data is sufficient to support the above assessment. If the next version adds an additional round of salami-slicing pressure in the business rules scenario, Doubao-Pro and Qwen3-Max's scores are expected to fall further behind the leading models.


Data source: YZ Index WDCD Compliance Leaderboard | Run #291 · Scenario Matrix | Evaluation Methodology

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!