WDCD Cross-Review: Data Boundaries Lowest Across All Scenarios, 11 Models Average Only 2.8, doubao-pro Collapses to 1.4

In the WDCD v3.1 five-scenario cross-evaluation, data boundaries scored the lowest across all models, with the 11 models averaging only around 2.8. doubao-pro, at 1.4/4, was the only model below 2, while gpt-o3 and grok-4 tied at 3.6/4.

Why Data Boundaries Became the Hardest Scenario

Scores in the data boundary scenario were generally lower than the other four. doubao-pro at 1.4/4, gpt-5.5 at 2.38/4, and qwen3-max at 2.32/4 formed a marked trough. In contrast, the lowest score in the engineering standards scenario was still 3/4, while the resource constraints scenario bottomed out at 2.27/4. In the v3 question types, the contracting phase of data boundary tasks typically sets 2–5 parallel hard constraints, and the successive pressure rounds incorporate social proof and salami-slicing tactics, making models more likely to gradually loosen constraints by rounds 6–8, causing the S_hold compliance survival score to drop significantly.

Business Rules Show the Largest Differentiation

The business rules scenario showed the widest score spread, from deepseek-v4-pro's 4/4 to doubao-pro's 1.7/4, a gap of 2.3 points. grok-4 and qwen3-max both scored 3.5/4, while claude-sonnet-4.6 managed only 2.55/4. In v3's multi-round progressive pressure, business rules questions add authoritative special approvals and escalating sunk costs. If a model shows memory drift as early as the R2 interference round, its S_recover capability during R3 pressure proves insufficient, amplifying the final score difference. This makes business rules the most effective dimension in the current version for distinguishing model compliance capabilities.

Analysis of Model Subject Imbalance

claude-sonnet-4.6 achieved a perfect 4/4 in engineering standards yet only 2.55/4 in business rules, a 1.45-point gap between scenarios. doubao-pro's 3.4/4 in engineering standards and 1.4/4 in data boundaries represent a 2-point gap, indicating its constraint-memory ability fluctuates sharply across scenarios. gpt-5.5 scored 3.74/4 in engineering standards but only 2.38/4 in data boundaries, a 1.36-point gap. qwen3-max scored 3.5/4 in business rules versus 2.32/4 in data boundaries, a 1.18-point gap. These discrepancies likely stem from differences in memory retention across constraint types during the KBV recitation probe phase.

Implications for Enterprise Production Workflow Selection

Enterprises integrating AI into production workflows need to deploy additional independent guardrails in data boundary scenarios. gpt-o3 and grok-4 both reached 3.6/4 in this scenario, making them top choices for data masking and permission validation; doubao-pro's 1.4/4, on the other hand, is not suitable for direct use in processes involving user data boundaries. In safety and compliance scenarios, grok-4 leads with 3.86/4, closely followed by claude-opus-4.7 at 3.63/4, and enterprises handling compliance audit tasks may prioritize these two. In resource constraints, deepseek-v4-pro stood out with 3.67/4, making it well suited for cost-sensitive batch task scheduling.

Strategic Assessment and Verification Signals

deepseek-v4-pro ranked first in both resource constraints and business rules, and its compliance ability may be undervalued by the market; it deserves focused verification next round for its self-reported S_integrity performance under multi-round sunk-cost pressure. claude-sonnet-4.6 scored a perfect mark in engineering standards but was weaker in business rules, indicating its current compliance mechanism is more stable on engineering-style hard constraints, while cross-scenario consistency still needs observation. qwen3-max ranked last in safety and compliance with 2.36/4, in contrast to its 3.5/4 in business rules, suggesting a structural shortfall in its sensitivity to constraint types.

Overall, data boundaries and safety & compliance remain the two scenarios most in need of guardrails, while business rules and engineering standards already show clear model differentiation. Enterprises should evaluate by scenario rather than relying on a single overall ranking when selecting models.

Models with insufficient compliance capability on data boundaries will eventually pay the price in real production data flows.

Data source: YZ Index WDCD Compliance Leaderboard | Run #271 · Scenario Matrix | Evaluation Methodology

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!