In the WDCD v3.1 five-scenario cross-review, the business rules scenario received the lowest scores across the board: Doubao-Pro scored only 2.25/4, and among the remaining models the highest was only Claude Opus 4.7 at 4/4. Overall performance was clearly weaker than in the other four constraint categories.
Why Business Rules Is the Hardest Scenario
The business rules scenario requires the model to remember multiple parallel hard constraints during the commitment phase and to maintain compliance under four subsequent rounds of pressure: social proof, authoritative special approval, salami slicing, and sunk cost. Doubao-Pro scored only 2.25/4 in this scenario, far below its 3.9/4 performance on data boundaries, indicating a significant drop in its S_hold score under multi-constraint parallel memory and continuous pressure. By contrast, Qwen3-Max also ranked last in the safety compliance scenario at 2.25/4, but more models fell below 3/4 in the business rules scenario (Claude Sonnet 4.6 at 2.8/4, GPT-5.5 at 3.25/4), suggesting that this scenario's KBV restatement probe and final-round honest self-report place higher demands on the model's constraint memory and self-consistency.
The Most Discriminative Scenario and Model Imbalances
The data boundaries scenario showed the greatest discrimination: GLM-4.6 scored only 1.7/4, while DeepSeek V4 Pro and Gemini 3.1 Pro both reached 4/4, a gap of 2.3 points. GLM-4.6's 1.7/4 on data boundaries and 3.4/4 on business rules form a 1.7-point gap, indicating that its S_kbv and S_recover capabilities are weakest under data-related hard constraints. Doubao-Pro's data boundaries score of 3.9/4 and business rules score of 2.25/4 differ by 1.65 points, while Claude Sonnet 4.6's safety compliance score of 4/4 and business rules score of 2.8/4 differ by 1.2 points, both showing clear imbalance.
Claude Opus 4.7 scored 4/4 in all three scenarios of resource limits, business rules, and safety compliance, yet only 3/4 in engineering norms, making it the only top-tier model to score lower than Gemini 2.5 Pro on engineering norms. This 1-point gap may stem from the engineering norms scenario placing greater emphasis on S_hold survival during the R3 pressure round; after Claude Opus 4.7 broke under that round, its S_recover recovery was insufficient.
Selection Implications for Enterprises Integrating AI into Production Processes
For enterprises integrating AI into production processes, additional guardrails must be deployed for Doubao-Pro, Claude Sonnet 4.6, and Qwen3-Max in the business rules scenario, because these models have the lowest commitment-survival rates under multi-round salami-slicing pressure. In the data boundaries scenario, GLM-4.6 should not be used alone; pairing it with DeepSeek V4 Pro or Gemini 3.1 Pro is recommended to reduce data leakage risk. In the engineering norms scenario, Claude Opus 4.7's 3/4 performance means it requires human review in hard-rule scenarios such as code review and version constraints, whereas Gemini 2.5 Pro and Gemini 3.1 Pro, both at 4/4, are better suited for direct embedding into CI/CD pipelines.
In the resource limits scenario, Doubao-Pro and Qwen3-Max both ranked last at 2.8/4. If enterprises rely on API quota controls, they should prioritize Claude Opus 4.7 or DeepSeek V4 Pro.
Strategic Judgment and Signals for Next-Round Verification
Claude Opus 4.7's commitment-keeping ability may be overestimated by the market given its full scores in four scenarios, and its 3/4 weakness in engineering norms could be amplified in real multi-turn conversations; GLM-4.6's 1.7/4 on data boundaries suggests its commitment-keeping ability is underestimated, as it exposes a weakness in only a single scenario. Doubao-Pro's 1.65-point imbalance gap deserves focused verification in the next round, especially in the sunk-cost pressure round of the v3 question, to see whether its S_integrity honest self-report score is 0.
In the safety compliance scenario, both Claude Sonnet 4.6 and Claude Opus 4.7 scored 4/4, indicating that the Claude series has the most stable combination of S_hold and S_kbv under this constraint type, which can serve as a baseline model for compliance-priority scenarios.
The low scores in the business rules scenario reveal that multi-constraint parallel memory remains the most fragile link in current models. If the next version achieves a breakthrough here, the overall commitment-keeping score ranking may be rewritten.
Data source: YZ Index WDCD Commitment-Keeping Leaderboard | Run #326 · Scenario Matrix | Evaluation Methodology
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接