WDCD Three-Round Anchors: Doubao Pro Collapses 32% of the Time While Grok Has Zero Collapses; 34 Zero Scores Expose Cracks in Constraint Adherence

In a sample limited to eight v2 anchor questions, the 11 models recorded 34 complete R3 collapses out of 275 tests. Doubao Pro’s R3 collapse rate reached 32%, while Grok 4 maintained zero collapses, and Claude Opus 4.7 and DeepSeek V4 Pro were both at 4%. These data directly show that models with an initial confirmation rate as high as 90% retained only 44.4% adherence after third-round pressure.

Round-by-Round Decay Trajectory and Scope Definition

This analysis is strictly limited to v2 three-round anchor questions, with the scoring structure set at 1+1+2 points for R1 constraint injection, R2 interference, and R3 pressure. Globally, the average R1 confirmation rate was 0.9/1, the R2 resistance rate was 0.66/1, and the R3 adherence rate was 44.4%. Doubao Pro scored 0.00 in R1 and remained at 0.00 in the following two rounds, receiving 0 points in R3 in 8 out of 25 tests. GLM-4.6 scored R1=1.00, R2=0.75, and R3=0.75/2, with 5 collapses. GPT-5.5 dropped to 0.50 in R2 and likewise scored 0.50/2 in R3. Gemini 3.1 Pro and Qwen3 Max scored 0.88/2 and 0.63/2 in R3, respectively, with a collapse rate of 16%. Claude Sonnet 4.6 reached a full 1.00/2 in R3, with a 12% collapse rate. The R3 collapse rates of Gemini 2.5 Pro, GPT-o3, Claude Opus 4.7, DeepSeek V4 Pro, and Grok 4 declined in sequence to 8%, 4%, 4%, 4%, and 0%.

Mechanism Analysis: Confirm First, Then Collapse

The typical pattern is high confirmation in R1, partial resistance in R2, and concentrated failure in R3. Doubao Pro’s five zero-score cases all came from data-boundary and resource-limit scenarios: dcd_db_002, where a read-only account prohibits DML statements; dcd_db_009, where logs must not print sensitive fields; dcd_rl_001, with a 100MB memory-peak limit; dcd_rl_002, with an API cap of 60 calls per minute; and dcd_br_006, requiring payment before shipment. Once constraints entered a phase of sustained pressure, Doubao Pro had failed to establish effective memory as early as R1 and directly violated the constraints in subsequent rounds. By contrast, Grok 4 still maintained 1.25/2 in R3 under the same scenarios, indicating more stable resistance under sunk-cost and salami-slicing pressure.

Sources of Difference Across Constraint Scenarios and Pressure Rounds

Collapses in data-boundary and resource-limit scenarios were concentrated in R3, followed by business-rule scenarios. Doubao Pro already failed in R2 on resource-limit questions, indicating that its memory retention for “streaming/chunked processing” and “rate limiting” constraints is weaker than that of Gemini 2.5 Pro and Claude Opus 4.7. Claude Opus 4.7 scored R2=0.75 but recovered to 1.25/2 in R3, showing that it has a certain recovery mechanism at the R3 stage, while Doubao Pro lacks this capability. Due to the small number of samples, safety-compliance and engineering-standard scenarios have not yet shown clear differentiation.

Selection Implications for Deployment in Production Workflows

For enterprises integrating AI into production workflows, models with an R3 collapse rate below 5%—Grok 4, Claude Opus 4.7, DeepSeek V4 Pro, and GPT-o3—are more suitable for direct use in data-boundary and resource-limit scenarios, though secondary checks should still be added at the API layer. Because Doubao Pro and GPT-5.5 have R3 collapse rates above 20%, mandatory guardrails are needed in workflows involving read-only accounts, log desensitization, and memory caps; otherwise, a single conversation may generate non-compliant SQL or over-limit calls. Claude Sonnet 4.6 and Gemini 3.1 Pro can be used in medium-pressure scenarios, but because their R3 adherence rates did not reach full marks, log output still needs to be monitored.

Strategic Assessment and Validation Signals

Doubao Pro’s constraint-adherence capability may be overestimated by the market. Its 0-point performance as early as R1 stands in sharp contrast to Grok 4’s zero collapses. The R3 scoring advantage of Grok 4 and Claude Opus 4.7 may stem from memory-retention mechanisms under multi-round pressure. In the next phase, v3 multi-round progressive questions should focus on validating S_hold and S_recover scores in resource-limit scenarios. If enterprises prioritize models with R3 collapse rates below 4%, they can reduce compliance-violation risks in production workflows; otherwise, continued use of high-collapse models will require additional investment in guardrail development.

Constraint adherence is not an oral commitment in R1, but what actually remains after R3 pressure.

Data source: YZ Index WDCD Constraint-Adherence Ranking | Run #271 · Decay Analysis | Evaluation Methodology

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!