In the WDCD v3.1 compliance test, the business rules scenario recorded the lowest average score, with Claude-sonnet-4.6 scoring only 1.8/4 and Grok-4 achieving a perfect 4/4, a gap of 2.2 points.
Business rules scenario becomes the hardest for all models
The business rules scenario posted the largest spread, with the lowest score of 1.8/4 and the highest of 4/4—far exceeding the other four scenarios. In the data boundary scenario, the lowest score was qwen3-max's 2.4/4; in resource constraints, the lowest was claude-sonnet-4.6's 2.33/4; in safety and compliance, the lowest was qwen3-max's 1.8/4; and in engineering standards, the lowest was gemini-2.5-pro's 3.18/4. The stress design of the business rules scenario includes successive rounds of "salami slicing" combined with "authoritative special approval," a constraint combination most likely to be gradually breached after the third round, causing a sharp drop in S_hold compliance survival scores.
The most discriminating scenario also points to business rules
The business rules scenario showed the largest standard deviation in scores, with a clear spread from Grok-4's 4/4 and claude-opus-4.7's 3.7/4 to Doubao-pro's 2.3/4, gemini-3.1-pro's 2.25/4, and claude-sonnet-4.6's 1.8/4. In contrast, in the engineering standards scenario, all models scored above 3.18/4, with the least discrimination. Although the safety and compliance scenario had a low minimum, the median models clustered in the 3.0–3.4 range, making it less discriminating than business rules.
Uneven model performance concentrated in three scenarios
Claude-sonnet-4.6 scored 3.7/4 on engineering standards but only 1.8/4 on business rules, a gap of 1.9 points. Gemini-3.1-pro scored 3.76/4 on engineering standards and 2.25/4 on business rules, a gap of 1.51 points. Qwen3-max scored 3.28/4 on engineering standards and 1.8/4 on safety and compliance, a gap of 1.48 points. Gpt-o3 scored 4/4 on engineering standards and 2.68/4 on data boundary, a gap of 1.32 points. These gaps emerged mainly during the continuous pressure rounds 6–9 of the v3 test, indicating structural differences in models' ability to retain constraints between "hard business process constraints" and "engineering code standards."
Resource constraints scenario reveals widespread weakness outside Grok-4
In the resource constraints scenario, claude-opus-4.7 scored only 2.83/4, gemini-2.5-pro also 2.83/4, and claude-sonnet-4.6 bottomed at 2.33/4. Grok-4 led at 3.75/4, followed by gemini-3.1-pro at 3.47/4 and deepseek-v4-pro at 3.45/4. The v3 test in this scenario emphasizes "sunk cost" pressure, causing models to accept additional resource requests after round 7, leading to loss of S_recover breach recovery scores.
Safety and compliance scenario: low-scoring models need reinforced guardrails
In the safety and compliance scenario, qwen3-max scored only 1.8/4, gpt-5.5 scored 2.54/4, and gemini-2.5-pro scored 2.56/4. Grok-4 scored 3.86/4, with deepseek-v4-pro and gpt-o3 tied at 3.43/4. The KBV recall probe in the safety and compliance scenario is triggered at round 10; qwen3-max repeatedly failed to accurately recall the initial compliance constraints in that round, causing its S_kbv constraint memory score to drop directly to zero.
Engineering standards scenario most balanced, GLM-4.6 and GPT-o3 achieve perfect scores
In the engineering standards scenario, glm-4.6 and gpt-o3 both scored 4/4, gemini-3.1-pro scored 3.76/4, grok-4 scored 3.74/4, and the remaining models all scored above 3.44/4. The v2 anchor questions in this scenario generally yielded high scores across three rounds, indicating that models' resistance to engineering constraints such as code style, version control, and test coverage is generally stronger than to business process constraints.
Scenario-based model selection recommendations for enterprise production integration
Enterprises integrating AI into business-rule-intensive processes should prioritize Grok-4 or claude-opus-4.7, and add independent rule-checking nodes at the process layer to avoid relying solely on the model's compliance. In resource-constrained scenarios, Grok-4, gemini-3.1-pro, and deepseek-v4-pro all score above 3.4/4 and can be first choices; claude-sonnet-4.6 scores only 2.33/4 in this scenario, requiring additional resource quota monitoring. In safety and compliance scenarios, qwen3-max and gpt-5.5 score below 2.6/4; an external compliance review layer must be deployed before integration. In engineering standards scenarios, most models can be used directly, with glm-4.6 and gpt-o3 as default options.
Strategic assessment: Grok-4's compliance capability is underestimated by the market
Grok-4 ranks first in business rules, resource constraints, and safety and compliance, and also enters the top four in engineering standards, making it the most balanced in overall compliance capability. Claude-sonnet-4.6 scores 3.7/4 on engineering standards but only 1.8/4 on business rules, suggesting its compliance capability is overestimated by part of the market. Qwen3-max bottoms out in both data boundary and safety and compliance scenarios; the next test should closely observe its KBV recall performance in rounds 8–10 of the v3 test. Glm-4.6 achieves a perfect score in engineering standards but only 3/4 in business rules; it is recommended to add a double-check mechanism in business process scenarios.
The 2.2-point gap in the business rules scenario indicates that unless next-generation models can maintain initial constraints under successive rounds of "salami slicing" pressure, rule-based applications in production will still rely on external guardrails.
Data source: YZ Index WDCD Compliance Leaderboard | Run #247 · Scenario Matrix | Evaluation Methodology
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接