Across the five scenario tests in WDCD v3.1, the business rules scenario was the lowest-scoring area for all models; doubao-pro scored only 2.13/4, far below the floor in other scenarios.
Why the Business Rules Scenario Is Hardest for Adherence
In the business rules scenario, the top group—deepseek-v4-pro, gemini-2.5-pro, gpt-5.5, and others—held steady at 3.67/4, but bottom-ranked doubao-pro scored only 2.13/4, and qwen3-max managed only 2.37/4. This gap stems from the sequential pressure stages in the v3 tasks: after a model first accepts the commitment that "business processes must be executed strictly according to rules," the test injects social proof, authoritative special approval, and salami-slicing exception requests in turn. Low-scoring models caved under the third round of salami-slicing pressure and failed to maintain the initial constraint.
By contrast, the engineering standards scenario had the highest overall scores: gpt-6-luna, gpt-6-sol, gpt-6.1-sol, gpt-o3, and grok-4 all received a perfect 4/4, while the lowest scorer, gemini-2.5-pro, still reached 3.31/4. Engineering standards constraints are mostly explicit, verifiable code or process rules, so models find it easier to maintain consistency during the KBV restatement probe stage.
The Two Most Discriminative Scenarios
The data boundary scenario had the highest discrimination: deepseek-v4-pro and grok-4 tied at 3.8/4, while gpt-6-sol scored only 2.16/4, a spread of 1.64 points. In this scenario, the v3 multi-turn pressure focuses on the hard constraint that "data boundaries must not be crossed," and the bottom-ranked model was the first to give in during the sunk-cost escalation stage. The discriminative power of the engineering standards scenario is reflected in specialization imbalance: most models score prominently there but drop quickly in other scenarios.
Model Imbalance and Its Causes
In engineering standards, deepseek-v4-pro scored 3.84/4, but in resource constraints it scored only 2.79/4, a gap of 1.05 points. The pressure path in the resource constraints scenario relies more on a model's long-term memory of "parallel hard constraints"; deepseek-v4-pro may have lost some constraints after the R2 interference round, causing its S_hold score to decline.
GPT-6-sol scored 4/4 in engineering standards but only 2.16/4 in data boundaries, a gap of 1.84 points. The model also scored only 3.11/4 in safety compliance, indicating that its adherence capability for "data must not be leaked" constraints is significantly weaker than its adherence capability for engineering processes. A similar imbalance is evident in gpt-6-luna: 4/4 in engineering standards and only 2.46/4 in data boundaries.
doubao-pro's imbalance is the most extreme: 3.41/4 in engineering standards and only 2.13/4 in business rules, a gap of 1.28 points. This indicates that its recovery capability (S_recover) when handling "business process exception authorization" is insufficient.
Model Selection Implications for Enterprises Integrating AI into Production Workflows
For enterprises integrating AI into production workflows, models scoring below 3.0/4 in the business rules scenario need an external rules engine added. The performance of doubao-pro and qwen3-max in this scenario means that directly opening business approval workflows could create a risk of rules being gradually eroded. Models scoring below 2.5/4 in the data boundary scenario (gpt-6-sol, gpt-6-luna, gpt-o3) must have a secondary human review checkpoint in scenarios involving user data or internal documents.
Models that score high in the engineering standards scenario (the gpt-6 series, grok-4) are suitable for code review, deployment pipelines, and similar stages, but enterprises still need to validate their performance separately in resource constraints scenarios, because deepseek-v4-pro's low score there shows that strong engineering capability does not equal strong resource quota control capability.
Strategic Assessment
Current data indicates that the GPT-6 series' perfect performance in engineering standards may be overestimated by the market; its systematic imbalance in data boundaries and safety compliance (with gaps in both exceeding 1.3 points) has not been fully priced in. Conversely, grok-4 scored 3.86/4 in safety compliance and tied for first in data boundaries, so its overall adherence capability may be underestimated. The next round of testing should focus on the specific failure point of low-scoring models in the business rules scenario under the R3 pressure round, to determine whether it is memory decay or deliberate compromise.
Low scores in the business rules scenario are not a question of model intelligence, but of how well constraints survive under multi-turn social pressure.
Data source: YZ Index WDCD Adherence Leaderboard | Run #360 · Scenario Matrix | Evaluation Methodology
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接