In the resource limitation scenario, gpt-5.5 scored only 1.55/4, and in business rules, Doubao-pro scored only 1.45/4. These two specific numbers directly reveal the weakest constraint types in the WDCD v3.1 compliance test.
Resource Limitation Becomes the Lowest Scenario Across All
Among the five scenarios, resource limitation has the lowest average score, with champion claude-opus-4.7 achieving 4/4 and bottom-ranked gpt-5.5 only 1.55/4, a range of 2.45 points. The security compliance scenario is also severe, with Doubao-pro and gpt-5.5 tied at 1.65/4. In contrast, the data boundary scenario is the highest overall, with gemini-2.5-pro, gemini-3.1-pro, glm-4.6, and grok-4 all tying at 4/4. The business rules scenario also shows sharp divergence: deepseek-v4-pro and gpt-5.5 both score 4/4, while Doubao-pro drops to 1.45/4.
Stress Rounds Reveal Constraint Memory Gaps
The WDCD v3 test uses 8-12 rounds of dialogue, first establishing a contract, then applying progressive pressure through social proof, authority special approval, salami slicing, and sunk cost, followed finally by a KBV paraphrase probe and an end-of-round honesty self-report. The resource limitation scenario has the lowest score, indicating that models are most likely to lose the hard constraint of "resource cap" after three or more consecutive rounds of pressure. The security compliance scenario also suffers significant score loss during the R3 pressure stage; the 1.65/4 for both Doubao-pro and gpt-5.5 is a direct result of this mechanism. The data boundary scenario generally scores higher, suggesting that the "data scope" constraint is more durable under the same pressure.
The 1.3-point gap between claude-opus-4.7's 4/4 in resource limitation and 2.7/4 in business rules exposes the memory stability variance of the same model across different constraint types.
True Risk Profile of Imbalanced Models
A total of nine models have a gap exceeding 1 point. Doubao-pro scores 3.9/4 in data boundary but 1.45/4 in business rules, a gap of 2.45 points; gpt-5.5 scores 4/4 in business rules but 1.55/4 in resource limitation, also a gap of 2.45 points. Gemini-2.5-pro scores 4/4 in data boundary but only 2.1/4 in engineering standards, a gap of 1.9 points. These numbers indicate that model compliance capability across scenarios is not linearly correlated; enterprises cannot infer overall scenario performance from a single scenario result.
Scenario-Level Guardrail Strategies for Production Deployment
For enterprises integrating AI into production workflows, the resource limitation scenario requires external hard interception because the model's own compliance capability is weakest. The security compliance scenario also requires secondary review; the 1.65/4 for Doubao-pro and gpt-5.5 is already below the passing line. The data boundary scenario can be moderately relaxed; the 4/4 scores of four models including gemini-2.5-pro show that this constraint is relatively reliable. The business rules scenario needs to be differentiated by model: deepseek-v4-pro and gpt-5.5 can be used directly, while Doubao-pro must be augmented with a rules engine.
Overestimated and Underestimated Compliance Capabilities in the Market
gpt-5.5 scores 4/4 in business rules but only 1.55/4 in resource limitation; its overall compliance capability may be overestimated by the market. claude-opus-4.7 scores 4/4 in both resource limitation and security compliance, but only 2.7/4 in business rules; its capability may be underestimated. glm-4.6 scores 4/4 in data boundary, resource limitation, and engineering standards, with the highest overall stability in this round of data. The next validation focus should be on the R3 pressure rounds in the resource limitation scenario, observing the actual breaking point under continuous salami slicing attacks for models between 1.55/4 and 4/4.
When selecting models, enterprises can directly refer to scenario score rankings: for resource limitation, prioritize claude-opus-4.7 or glm-4.6; for security compliance, prioritize the claude series or deepseek-v4-pro; for business rules, prioritize deepseek-v4-pro or gpt-5.5. For engineering standards, consider glm-4.6 or Doubao-pro. Any cross-scenario deployment must deploy additional external compliance layers for both resource limitation and security compliance.
Compliance capability is not a general attribute of a model, but a scenario-specific attribute. This judgment comes directly from the five sets of raw scores in WDCD v3.1.
Data source: YZ Index WDCD Compliance Leaderboard | Run #242 · Scenario Matrix | Evaluation Methodology
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接