In the WDCD v3.1 five-scenario comparative evaluation, the Business Rules scenario became the track with the lowest scores across all models, with qwen3-max managing only 1.55/4, far below gemini-3.1-pro's 3.85/4.
Why Business Rules Became the Hardest Scenario
The champion of the Business Rules scenario, gemini-3.1-pro, only reached 3.85/4, while the last-place qwen3-max fell to 1.55/4, a spread of more than 2.3 points. By comparison, four models simultaneously achieved a perfect 4/4 in the Data Boundary scenario, and three models scored full marks in the Safety & Compliance scenario. The pressure design of the Business Rules scenario includes consecutive multi-round social approval and salami-slicing tactics, which may make models more prone to gradually loosening their stance under parallel hard constraints. qwen3-max's low score in this scenario directly dragged down its overall compliance performance.
The Second Layer of Pressure in the Resource Constraint Scenario
The Resource Constraint scenario is equally difficult; claude-opus-4.7 scored only 2/4, ranking last in this scenario. glm-4.6 led with 4/4, while gemini-3.1-pro came in at 3.7/4. claude-opus-4.7 achieved 4/4 in Data Boundary but dropped to 2/4 in Resource Constraints, a 2-point gap between scenarios for a single model. This shows that different types of constraints impose markedly different compliance requirements on the same model. The v3 multi-round progressive pressure in the Resource Constraint scenario may trigger the sunk cost effect earlier, making it easier for the model to break down at the R3 stage.
Actual Distribution of Lopsided Model Performance
Eight models showed lopsided score gaps exceeding 1 point across scenarios. claude-opus-4.7 formed the largest gap between its 4/4 in Data Boundary and 2/4 in Resource Constraints; doubao-pro scored 4/4 in Data Boundary but only 2/4 in Business Rules, also a 2-point gap; qwen3-max's gap between 3.7/4 in Data Boundary and 1.55/4 in Business Rules reached 2.15 points. gpt-5.5 scored 4/4 in Engineering Standards but only 2.4/4 in Safety & Compliance, a 1.6-point gap. These figures indicate that no single model can maintain a balanced high level across all five constraint scenarios.
Selection Implications for Enterprises Integrating AI into Production Workflows
Enterprises integrating AI into production workflows need to match models to their specific scenarios. In the Data Boundary and Safety & Compliance scenarios, both claude-opus-4.7 and grok-4 can consistently achieve 4/4, making them priority candidates for sensitive data processing and compliance review. For the Resource Constraint scenario, claude-opus-4.7 should be avoided in favor of glm-4.6 or gemini-3.1-pro. In the Business Rules scenario, gemini-3.1-pro's 3.85/4 performance is relatively stable, while qwen3-max's 1.55/4 means this model requires additional guardrails or secondary verification under such constraints.
Five models achieved 4/4 in the Engineering Standards scenario, including claude-sonnet-4.6, doubao-pro, and glm-4.6, making them suitable for code review and process control stages. If an enterprise spans multiple scenarios, a model routing mechanism should be in place to avoid exposing risk by relying on a single model in its weakest scenario.
Strategic Assessment and Next-Round Verification Signals
claude-opus-4.7's perfect scores in Data Boundary and Safety & Compliance may lead the market to overestimate it as reliable across all scenarios, yet its actual performance of 2/4 in Resource Constraints reveals a clear shortfall in compliance capability. qwen3-max's extremely low score in Business Rules may also cause its overall compliance capability to be underestimated. glm-4.6 achieved 4/4 in both Resource Constraints and Engineering Standards, making it worth continued observation in the next round for its recovery capability under multi-round pressure.
This pilot sampling adopted a worst-of-3 methodology. If KBV recitation probe rounds are added to v3 questions in the future, changes in the S_kbv and S_recover scores of lopsided models will become important verification signals. When selecting models, enterprises should treat scenario-level score differences as hard inputs rather than relying solely on a single composite ranking.
No model can achieve perfect scores across all constraint scenarios at once; model selection is essentially about allocating risk to the best-matched model.
Data source: YZ Index WDCD Compliance Leaderboard | Run #276 · Scenario Matrix | Evaluation Methodology
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接