This period's WDCD v3.1 testing shows that Qwen3 Max fell 21.9 points from Run #360, Gemini 2.5 Pro fell 10.4 points, Claude Sonnet 4.6 fell 7.1 points, DeepSeek V4 Pro fell 6.4 points, GPT-5.5 fell 5.6 points, and Grok 4 fell 5.2 points. A total of six models showed negative changes, while only Doubao Pro achieved a 9.5-point rise.
Data Facts: Declines Concentrated in Continuous Pressure Stage
The WDCD v3 question type uses 8–12 rounds of dialogue: it first establishes 2–5 hard constraints, then applies social proof, authority special approval, salami-slicing, and sunk-cost pressure in sequence, and finally scores via a KBV recall probe and a final-round honest debrief. S_hold constraint-keeping survival has a weight of 60 points; the later the breach, the higher the score. In this period's worst-of-3 sampling, Qwen3 Max most likely breached simultaneously during the salami-slicing and sunk-cost stages in rounds 5–7, causing a large loss in its S_hold score. Gemini 2.5 Pro's 10.4-point drop likewise points to middle-to-late pressure rounds rather than the initial constraint-setting stage.
Doubao Pro, however, achieved a 9.5-point improvement on the same question pool, indicating that its S_recover breach-recovery capability under multiple rounds of gradual pressure may have strengthened. In the current Top 5, GLM-4.6 reached 95.20 points, Grok 4 scored 94.80 points, Claude Opus 4.7 scored 94.00 points, GPT-o3 scored 93.50 points, and Gemini 3.1 Pro scored 93.20 points. Although Grok 4 fell 5.2 points, it still holds second place, indicating that its baseline constraint-keeping level remains higher than most participating models.
Cause Analysis: Differences in Constraint Scenarios and Pressure Rounds
The five types of constraint scenarios include data boundaries, resource limits, business rules, security compliance, and engineering standards. The v3 questions use authority special approval and sunk-cost pressure and are most likely to expose vulnerability in security compliance and business rule scenarios. Qwen3 Max fell 21.9 points, far more than other models; it is inferred that under security compliance constraints, it has the weakest resistance to authority special approval, and its breach point occurred significantly earlier. Gemini 2.5 Pro's 10.4-point drop may be concentrated in salami-slicing pressure in resource-limit scenarios, failing to maintain initial constraints after three consecutive rounds of interference.
Doubao Pro's 9.5-point rise may come from improvement in the S_integrity honest self-report 15-point item, meaning it less often falsely claims innocence after a breach. The v2 anchor question's three-round design (R1 injects constraints, R2 interferes, R3 applies pressure) has a maximum score of 4 points. This period's equal-weighted average calculation shows that most declining models lost heavily in the R3 pressure stage, while Doubao Pro scored relatively stably in that stage.
Selection Implications: Real Risks of Integration into Production Processes
For enterprises integrating AI into production processes, WDCD scores directly correspond to guardrail requirements in different scenarios. GLM-4.6 and Grok 4 are in the 95-point range, suitable for directly handling business rule and engineering standard constraints without additional multi-layer prompt protection. Security compliance scenarios require caution: Qwen3 Max's performance this period suggests it is prone to breaching under authority pressure. If enterprises use this model for compliance review, they must add independent rule-engine verification.
After Gemini 2.5 Pro's 10.4-point drop, its reliability in resource-limit scenarios has decreased. It is recommended to set up a secondary confirmation mechanism in cost-control or quota-management tasks. Although Doubao Pro's rise is positive, it remains below the Top 5 average, making it suitable as a low-risk supporting role rather than a core decision node.
Strategic Judgment: Overestimated and Underestimated Signals
Qwen3 Max's sharp decline this period indicates that its constraint-keeping ability was overestimated by prior market expectations. The next period needs to verify whether it can recover its S_hold score in authority special approval rounds in security compliance scenarios. GLM-4.6 leads with 95.20 points, and combined with Grok 4's 94.80 points, this shows that open-source/domestic models have formed a local advantage in multi-round constraint retention, making them worth prioritizing in enterprise candidate pools.
Claude Opus 4.7 and GPT-o3 rank third and fourth with 94.00 and 93.50 points, respectively, indicating that closed-source flagship models still perform stably on S_kbv constraint memory and S_recover recovery, but their lead has narrowed to the 1–2 point range. In the next period, focus should be on observing changes in the breach timing of Qwen3 Max and Gemini 2.5 Pro in rounds 6–8 of the v3 questions to confirm whether the trend continues.
Constraint-keeping ability is no longer a static label but a dynamic survival capability under each round of pressure; this period's data has clearly drawn the boundary between high-risk models and usable models.
Data source: YZ Index WDCD Constraint-Keeping Leaderboard | Run #365 · Change Tracking | Evaluation Methodology
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接