AI News

Data Boundaries Emerge as the Biggest Compliance Blind Spot: 11 Models Score as Low as 1.3, with Gaps up to 2.7

In WDCD v3.1's five constraint scenario tests, data boundary scenarios had the lowest average scores, with Qwen3-Max scoring only 1.3/4 and GLM-4.6 scoring 1.7/4, in stark contrast to six models achieving 4/4 in business rules scenarios. The results reveal severe model specialization and significant differences in retaining different types of constraints.

WDCD Compliance Test 场景横评
29