WDCD Five-Scenario Cross-Evaluation: Business Rules Lowest Across All Models, Engineering Standards Show Cruel 3-Point Gap

Claude-opus-4.7 scored 4/4 in data boundaries, resource constraints, and security compliance, but only 3/4 in engineering standards, making it the most prominent case of uneven performance in the WDCD v3.1 five-scenario cross-evaluation.

Business Rules Become the Toughest Scenario for All Models

In the business rules scenario, the average score of 11 models was significantly lower than that of the other four categories. The top models — claude-opus-4.7, deepseek-v4-pro, gpt-o3, and grok-4 — each scored only 3.5/4, while the lowest, Doubao-pro, scored just 1.5/4, a gap of 2 points. The v3 question bank for this scenario employed 8–12 consecutive dialogue turns: first establishing 2–5 parallel hard business constraints, then applying four levels of pressure — social conformity, authoritative special approval, salami slicing, and sunk cost — and finally conducting a KBV recall probe. Lower-scoring models typically broke the agreement by rounds 5–7, suffering the greatest loss in the S_hold compliance survival metric.

Engineering Standards Show the Greatest Discernment

The score range in the engineering standards scenario reached 3 points: deepseek-v4-pro and GLM-4.6 scored 4/4, while gemini-2.5-pro scored only 1/4. The v3 questions in this scenario focused on testing hard engineering constraints such as code standards, interface contracts, and version compatibility, with pressure applied primarily through sunk cost and salami slicing. Gemini-2.5-pro accepted a "temporary solution" that violated the norms as early as the third round, and during the subsequent KBV recall, it could not accurately remember the initial constraints, resulting in losses in both S_kbv and S_recover.

The Real Cost of Uneven Model Performance

Gemini-2.5-pro scored 4/4 in data boundaries but only 1/4 in engineering standards, a gap of 3 points between scenarios. Similarly, Doubao-pro scored 3.9/4 in data boundaries but dropped to 1.5/4 in business rules, a gap of 2.4 points. If such models are deployed in production processes requiring strict engineering standards or business rules, they are highly likely to break under pressure in rounds 6–8.

Claude-opus-4.7 showed a smaller gap of only 1 point, but its S_integrity honest self-report score in the engineering standards scenario was still lower than in other scenarios. GPT-o3 scored 4/4 in data boundaries and 2.65/4 in engineering standards, a gap of 1.35 points, indicating that its long-term memory for code-level constraints is weaker than for data boundary constraints.

Implications for Enterprise Model Selection

Enterprises integrating AI into production processes must deploy additional guardrails in the business rules scenario. The four models — claude-opus-4.7, deepseek-v4-pro, gpt-o3, and grok-4 — all scored 3.5/4 in this scenario, still leaving a 0.5-point gap. It is recommended to add manual review nodes or secondary constraint injection for models scoring below 3.5/4.

For the engineering standards scenario, deepseek-v4-pro and GLM-4.6 are top picks, with both scoring 4/4 and showing stable performance under R3 pressure rounds in v2 anchor questions. Gemini-2.5-pro and gemini-3.1-pro scored 1/4 and 2.75/4 respectively in this scenario, and are not recommended for direct use in code generation or deployment pipelines.

In the security compliance scenario, Qwen3-max scored only 1.8/4 and Doubao-pro scored 1.6/4, both below gpt-o3's 3/4. If an enterprise is involved in compliance audits, these two models require a forced context reset after round 4.

Strategic Assessment

Claude-opus-4.7's compliance ability is overestimated by the market across four scenarios. Its 3/4 in engineering standards compared to deepseek-v4-pro's 4/4 may stem from differences in long-term memory mechanisms for code contract constraints, which warrants focused validation in the upcoming v3.2. Deepseek-v4-pro scored 4/4 in both engineering standards and security compliance, suggesting its current market positioning may be undervalued.

The extreme contrast between gemini series performance in data boundaries and engineering standards indicates that its context decay mechanism performs unevenly across different constraint types. This signal should be continuously monitored in the next sampling round, especially on the worst-of-3 path.

Models with low scores in both business rules and engineering standards will be the first to be eliminated in production deployment.

Data source: YZ Index WDCD Compliance Ranking | Run #253 · Scenario Matrix | Evaluation Methodology

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!