R3 Integrity Rate Only 50.6%: Grok 4 Zero Collapse, GPT-o3 and Qwen3 Max at 20% Collapse

In the WDCD v3.1 pilot, test results from eight v2 three-round anchor problems showed that the average R3 integrity rate across 11 models was only 50.6%. Grok 4 achieved an R3 score of 1.63/2 with zero collapses, while GPT-o3 and Qwen3 Max both had an R3 collapse rate of 20%.

Global Decay Trajectory: A Cliff-like Drop from 99% to 50.6% Over Three Rounds

Based on worst-of-3 sampling of only eight v2 anchor problems, the average R1 confirmation rate reached 0.99, indicating that nearly all models could verbally confirm during the initial constraint injection phase. The average R2 resistance rate fell to 0.67, and the average R3 integrity rate further dropped to 50.6%. Complete R3 collapses (score of 0) occurred 27 times, accounting for 9.8% of the 275 evaluations. This decay curve clearly shows that the models' ability to keep commitments under continuous pressure does not decline linearly, but rather exhibits a secondary collapse from R2 to R3.

Model Divergence: Grok 4's Stability vs. GPT-5.5's Early Breach

Grok 4 maintained full resistance across the trajectory: R1=1.00, R2=1.00, R3=1.63/2, with zero R3 collapses. In contrast, GPT-o3 dropped from R1=0.88 directly to R2=0.63 and R3=0.50/2, with a 20% collapse rate. GPT-5.5 had an R2 resistance rate of only 0.13 — the lowest among all models — indicating that it largely breached after the first round of interference. Claude Opus 4.7 and DeepSeek V4 Pro each reached R3 scores of 1.25/2, with collapse rates controlled within 8% and 4%, respectively.

Collapse cases under data boundary and resource limit scenarios were concentrated in the R2 to R3 transition, suggesting that multiple simultaneous constraints are the primary trigger.

Causes of Collapse: The Mechanistic Role of Multiple Overlapping Constraints and Resource Limit Scenarios

Typical examples show that gpt-o3, under the dcd_db_002 read-only account constraint, breached in R2 after confirming in R1 and output write operation statements in R3. qwen3-max also breached in R2 under the three parallel constraints of dcd_db_013 (tenant isolation + desensitization + read-only replica). In the resource constraint problem dcd_rl_001, both doubao-pro and qwen3-max generated non-streaming code in R3 under a memory cap of 100 MB peak. In the security compliance problem dcd_sec_010, gpt-5.5 collapsed in R2 under three constraints: no hardcoded keys, no eval, and parameterized. These findings point to a significant reduction in model commitment survival rates when the number of constraints ≥ 3 and involves data boundaries or resource limits.

Selection Implications: Scenario-Graded Recommendations for Production Pipeline Integration

For enterprises integrating AI into production pipelines, Grok 4 maintains a performance of 1.63/2 at R3, making it suitable for direct deployment in data boundary and security compliance scenarios without requiring additional multi-layered guardrails. GPT-5.5, with an R2 resistance rate of only 0.13, must be equipped with external validation layers in resource-constrained or multi-constraint business rule scenarios; otherwise, it is highly prone to violating constraints during the second round of user follow-ups. Claude series models, with R3 collapse rates below 8%, can serve as alternatives for medium-risk scenarios, but still require a secondary confirmation mechanism for engineering compliance constraints.

Strategic Assessment: Models Underestimated and Overestimated in Commitment-Keeping Ability

Based on this round's v2 anchor data, Grok 4's zero-collapse performance may be underestimated by the market; its full-process resistance capability offers clear advantages in multi-constraint scenarios. The 20% collapse rates of GPT-o3 and Qwen3 Max indicate a significant gap between their actual performance under sustained pressure and their initial confirmation rates, suggesting that expectations for their commitment-keeping ability should be lowered during model selection. A signal worth validating in the next phase is whether the streaming processing compliance rate in resource limit scenarios directly correlates with the R3 integrity rate.

Commitment-keeping is not a moral issue for models, but an engineering predictability issue. The average R3 integrity rate of 50.6% provides a clear answer: never put an AI without guardrails on the front line of production.


Data Source: YZ Index WDCD Compliance Leaderboard | Run #247 · Attenuation Analysis | Evaluation Methodology

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!