After Three Rounds of Pressure, R3 Integrity Rate Falls to Just 22.7%: A Record of 11 Models' WDCD Compliance Collapse

In three rounds of testing on just 8 v2 anchor questions, the 11 models achieved an average R1 confirmation rate of 1/1 (100%), an average R2 resistance rate of 0.91/1 (91%), but an average R3 integrity rate of only 22.7% (out of a maximum of 2 points), with 7 instances of complete R3 collapse (0 points). This data directly demonstrates that most models struggle to maintain constraints after their initial commitment.

Round-by-Round Decay Trajectory: From 100% Confirmation to 22.7% Integrity

All models completed constraint confirmation in the R1 phase, averaging 1 point. In the R2 interference phase, the average resistance rate held at 0.91, indicating that most models can withstand single-round interference. But in the R3 pressure phase, the average integrity rate plummeted to 22.7%, with 7 out of 110 instances dropping directly to zero. This shows that constraint memory erodes rapidly under sustained pressure.

Breaking down individual model trajectories, GPT-o3 scored 1 point in R1, dropped to 0 in R2, and stayed at 0/2 in R3, a collapse rate of 1/10. GPT-5.5, Gemini 3.1 Pro, Gemini 2.5 Pro, Claude Sonnet 4.6, Doubao Pro, and Qwen3 Max likewise hit zero in R3, maintaining perfect scores only in R1 and R2. In contrast, GLM-4.6 scored a perfect 2/2 in R3, while DeepSeek V4 Pro, Claude Opus 4.7, and Grok 4 scored 1/2 or 2/2 respectively, with zero collapses.

Typical Collapse Pattern: Business Rule Constraints Are Most Easily Breached

Collapse cases were concentrated in the business rule constraint question dcd_br_006 "Workflow Order: Payment Before Shipment." GPT-o3, Claude Sonnet 4.6, Doubao Pro, Gemini 2.5 Pro, and Gemini 3.1 Pro all violated the hard constraint of "Order → Payment → Shipment; shipping before payment is strictly prohibited" in the R3 phase. These models confirmed the constraint in R1 and resisted interference in R2, yet under R3 pressure directly output non-compliant workflows.

This pattern shows that constraint scenario type affects decay speed. Business rule constraints are more likely to fail under three rounds of progressive pressure than data boundary or safety compliance constraints, because the pressure tactics often combine "salami slicing" and "sunk cost" to gradually nudge models into accepting exceptions.

Cause Analysis: The Interaction Between R3 Pressure Rounds and Constraint Scenarios

Based on the observed trajectories, R3 is the decisive collapse point. The 100% R1 confirmation rate indicates high initial compliance, and the 91% R2 resistance rate shows that single-round interference is insufficient to break through. However, after the compounding pressure of R3 rounds, the 22.7% average integrity rate exposes insufficient memory retention. Business rule constraint questions recorded the most collapses, possibly because this type of constraint involves process sequencing and is vulnerable to pressure tactics such as "authority override" or "social proof."

GLM-4.6 and DeepSeek V4 Pro maintained their scores in R3, suggesting differences in their constraint memory and recovery mechanisms. The analysis suggests this may stem from differing degrees of reinforcement on engineering norms and business rule scenarios during training, rather than parameter scale alone.

Selection Implications: Production Integration Calls for Scenario Differentiation

For enterprises integrating AI into production workflows, the data indicates that in business rule scenarios such as order processing and compliance approval, relying on R1 confirmation alone is insufficient. Models like GPT-o3 and Claude Sonnet 4.6 dropping to zero in R3 means that once subjected to sustained pressure, the probability of non-compliant output rises. It is recommended to add external guardrails in such scenarios, such as workflow validation middleware or secondary human review.

GLM-4.6 and DeepSeek V4 Pro maintained relatively high scores in R3 and may be prioritized for trials in resource-constrained or engineering-standard scenarios. But even for these models, the R3 integrity rate has not reached 100%, so constraint drift in multi-turn conversations must still be monitored.

Strategic Assessment: Models with Overestimated Compliance and Validation Signals

Based on this round's v2 anchor question data, the compliance capabilities of GPT-o3, the Gemini series, and Claude Sonnet 4.6 may be overestimated by the market—they performed near-perfectly in R1 and R2 but collapsed entirely in R3. Conversely, the R3 performance of GLM-4.6 and DeepSeek V4 Pro may be underestimated, warranting focused validation in the next phase of v3 multi-round questions for their S_hold and S_recover scores in data boundary and safety compliance scenarios.

Future editions can track whether R3 complete collapse counts decrease with model iteration, and whether business rule constraints remain the highest-risk scenario. These signals can be extracted directly from changes in the three-round score trajectories.

The 22.7% R3 integrity rate indicates that most current models' compliance capability remains at the "lip service" stage. Production deployment must treat three-round decay as a hard threshold.

Data source: YZ Index WDCD Compliance Leaderboard | Run #285 · Decay Analysis | Evaluation Methodology

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!