In a sample limited to eight v2 anchor questions, the 11 models recorded 34 complete R3 collapses out of 275 tests. Doubao Pro’s R3 collapse rate reached 32%, while Grok 4 maintained zero collapses, and Claude Opus 4.7 and DeepSeek V4 Pro were both at 4%. These data directly show that models with an initial confirmation rate as high as 90% retained only 44.4% adherence after third-round pressure.
Round-by-Round Decay Trajectory and Scope Definition
This analysis is strictly limited to v2 three-round anchor questions, with the scoring structure set at 1+1+2 points for R1 constraint injection, R2 interference, and R3 pressure. Globally, the average R1 confirmation rate was 0.9/1, the R2 resistance rate was 0.66/1, and the R3 adherence rate was 44.4%. Doubao Pro scored 0.00 in R1 and remained at 0.00 in the following two rounds, receiving 0 points in R3 in 8 out of 25 tests. GLM-4.6 scored R1=1.00, R2=0.75, and R3=0.75/2, with 5 collapses. GPT-5.5 dropped to 0.50 in R2 and likewise scored 0.50/2 in R3. Gemini 3.1 Pro and Qwen3 Max scored 0.88/2 and 0.63/2 in R3, respectively, with a collapse rate of 16%. Claude Sonnet 4.6 reached a full 1.00/2 in R3, with a 12% collapse rate. The R3 collapse rates of Gemini 2.5 Pro, GPT-o3, Claude Opus 4.7, DeepSeek V4 Pro, and Grok 4 declined in sequence to 8%, 4%, 4%, 4%, and 0%.
Mechanism Analysis: Confirm First, Then Collapse
The typical pattern is high confirmation in R1, partial resistance in R2, and concentrated failure in R3. Doubao Pro’s five zero-score cases all came from data-boundary and resource-limit scenarios: dcd_db_002, where a read-only account prohibits DML statements; dcd_db_009, where logs must not print sensitive fields; dcd_rl_001, with a 100MB memory-peak limit; dcd_rl_002, with an API cap of 60 calls per minute; and dcd_br_006, requiring payment before shipment. Once constraints entered a phase of sustained pressure, Doubao Pro had failed to establish effective memory as early as R1 and directly violated the constraints in subsequent rounds. By contrast, Grok 4 still maintained 1.25/2 in R3 under the same scenarios, indicating more stable resistance under sunk-cost and salami-slicing pressure.
Sources of Difference Across Constraint Scenarios and Pressure Rounds
Collapses in data-boundary and resource-limit scenarios were concentrated in R3, followed by business-rule scenarios. Doubao Pro already failed in R2 on resource-limit questions, indicating that its memory retention for “streaming/chunked processing” and “rate limiting” constraints is weaker than that of Gemini 2.5 Pro and Claude Opus 4.7. Claude Opus 4.7 scored R2=0.75 but recovered to 1.25/2 in R3, showing that it has a certain recovery mechanism at the R3 stage, while Doubao Pro lacks this capability. Due to the small number of samples, safety-compliance and engineering-standard scenarios have not yet shown clear differentiation.
Selection Implications for Deployment in Production Workflows
For enterprises integrating AI into production workflows, models with an R3 collapse rate below 5%—Grok 4, Claude Opus 4.7, DeepSeek V4 Pro, and GPT-o3—are more suitable for direct use in data-boundary and resource-limit scenarios, though secondary checks should still be added at the API layer. Because Doubao Pro and GPT-5.5 have R3 collapse rates above 20%, mandatory guardrails are needed in workflows involving read-only accounts, log desensitization, and memory caps; otherwise, a single conversation may generate non-compliant SQL or over-limit calls. Claude Sonnet 4.6 and Gemini 3.1 Pro can be used in medium-pressure scenarios, but because their R3 adherence rates did not reach full marks, log output still needs to be monitored.
Strategic Assessment and Validation Signals
Doubao Pro’s constraint-adherence capability may be overestimated by the market. Its 0-point performance as early as R1 stands in sharp contrast to Grok 4’s zero collapses. The R3 scoring advantage of Grok 4 and Claude Opus 4.7 may stem from memory-retention mechanisms under multi-round pressure. In the next phase, v3 multi-round progressive questions should focus on validating S_hold and S_recover scores in resource-limit scenarios. If enterprises prioritize models with R3 collapse rates below 4%, they can reduce compliance-violation risks in production workflows; otherwise, continued use of high-collapse models will require additional investment in guardrail development.
Constraint adherence is not an oral commitment in R1, but what actually remains after R3 pressure.
Data source: YZ Index WDCD Constraint-Adherence Ranking | Run #271 · Decay Analysis | Evaluation Methodology
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接