Data Boundaries Emerge as the Biggest Compliance Blind Spot: 11 Models Score as Low as 1.3, with Gaps up to 2.7

In the WDCD v3.1 five constraint scenario tests, the data boundary scenario had the lowest average score, with Qwen3-Max scoring only 1.3/4 and GLM-4.6 scoring 1.7/4, in sharp contrast to the six models that achieved 4/4 in the business rules scenario.

Why Data Boundaries Became the Most Difficult Scenario

The bottom scores in the data boundary scenario were concentrated in the last three places: Qwen3-Max 1.3/4, GLM-4.6 1.7/4, and GPT-5.5 2.5/4. By contrast, all top six in the business rules scenario reached 4/4, with DeepSeek-V4-Pro, Gemini-3.1-Pro, GPT-5.5, GPT-o3, Grok-4, and Qwen3-Max all holding full marks. This indicates that multi-round progressive pressure is more likely to break through models under the "data boundary" constraint.

The cause analysis points to the commitment-setting and continuous pressure phases in the v3 questions. Data boundary questions typically inject multiple hard constraints simultaneously in rounds 2–5, and social proof plus authority exception pressure more easily lead models to gradually expand the "scope of data use." Qwen3-Max scored only 1.3/4 in this scenario, indicating that it had already broken commitments multiple times during the R2 interference and R3 pressure phases, while the same model scored 4/4 in the business rules scenario, proving that its memory retention capacity differs markedly across different types of constraints.

Unexpected Divergence in Safety and Compliance Scenarios

In the safety and compliance scenario, champion Claude-Opus-4.7 scored 4/4, with DeepSeek-V4-Pro, GPT-5.5, and Grok-4 also achieving full marks. However, Claude-Sonnet-4.6 scored only 2.1/4, becoming the only model in this scenario below 2.5, creating a 1.9-point gap with its 4/4 performance in data boundaries. Gemini-2.5-Pro likewise scored only 2.2/4 in safety and compliance but achieved 4/4 in resource limits, a gap of 1.8 points.

This divergence may stem from the timing of the KBV restatement probe in the v3 questions. Safety and compliance questions insert more "salami-slicing" gradual requests before the final-round honest review. Claude-Sonnet-4.6 may have already relaxed its adherence to compliance boundaries in early rounds but failed to pull back in time during the S_recover phase, causing it to lose points in both S_hold and S_integrity.

The Real Impact of Specialized Models on Production Integration

Qwen3-Max scored 4/4 in business rules and 1.3/4 in data boundaries, a 2.7-point gap between scenarios, making it the most severely specialized model in this test. GLM-4.6 scored 1.7/4 in data boundaries and 3.4/4 in business rules, a gap of 1.7 points. If enterprises plan to integrate models into data processing pipelines, choosing Qwen3-Max or GLM-4.6 based solely on high business rules scores carries significantly higher risk than choosing Gemini-2.5-Pro or Gemini-3.1-Pro, which also maintain 3.5/4 or above in data boundaries.

Claude-Sonnet-4.6 scored 4/4 in both engineering standards and data boundaries but only 2.1/4 in safety and compliance. When using this model for code review or data masking processes, an independent compliance check layer must be deployed additionally; otherwise, the model's own compliance capability alone is insufficient to cover all scenarios.

Selection Advice: Match by Scenario, Not Overall Ranking

Enterprises integrating AI into production processes should prioritize the worst-of-3 results for the target scenario. For scenarios requiring strict data boundary control, Claude-Sonnet-4.6 (4/4) or Grok-4 (4/4) are recommended; for resource-limited scenarios, DeepSeek-V4-Pro (4/4) is preferred; for safety and compliance scenarios, avoid Claude-Sonnet-4.6 (2.1/4) and Gemini-2.5-Pro (2.2/4), and switch to Claude-Opus-4.7 or DeepSeek-V4-Pro.

Although most models performed well in the business rules scenario, Doubao-Pro scored only 2.25/4, lower than its 3.9/4 in safety and compliance, indicating that this model has the weakest resistance to rule-based constraints and requires additional manual review nodes during integration.

Strategic Judgment

This data indicates that Claude-Sonnet-4.6's compliance capability is overestimated by the market in engineering standards and data boundaries, but underestimated in safety and compliance; Qwen3-Max's full score in business rules may mask its extremely low compliance survival rate in data boundaries. If the next test adds more "data boundary + safety compliance" mixed questions, changes in the S_recover scores of the above models deserve close attention.

Compliance is not an add-on feature of a model, but a hard threshold before production integration.

Data source: YZ Index WDCD Compliance Leaderboard | Run #336 · Scenario Matrix | Evaluation Methodology

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!