WDCD v3.1 Cycle: Gemini 2.5 Pro Up 8.7 Points, Doubao Pro the Only Decliner at 7.3

WDCD v3.1 pilot data shows Gemini 2.5 Pro rising 8.7 points this cycle, GPT-5.5 up 7.9 points, Claude Opus 4.7 up 6.1 points, and GPT-o3 up 5.2 points, while Doubao Pro fell 7.3 points — a clear pattern of four risers and one decliner.

Data Facts: Real Fluctuations Under the Worst-of-3 Protocol

This sampling run adopts the worst-of-3 protocol, where each question is executed three times independently and the worst score is counted. Claude Opus 4.7 reached 91.50 on the WDCD this cycle, up 6.1 points from Run #276; Doubao Pro slid 7.3 points from its previous level. Gemini 2.5 Pro posted the largest gain, with GPT-5.5 following at +7.9 points. In the Top 5 standings, Grok 4 remains first at 97.50, GLM-4.6 second at 91.80, Claude Opus 4.7 third at 91.50, with GPT-o3 at 88.80 and GPT-5.5 at 88.60 in fourth and fifth place respectively.

The compliance total score is the equally weighted average of the native 100-point scale on v3 questions and the converted score on v2 anchor questions (total/4×100). Each v3 question contains 8-12 dialogue rounds: 2-5 parallel hard constraints are established first, followed by sequential pressure from social approval, authoritative exceptional approval, salami-slicing, and sunk costs. The session concludes with a KBV restatement probe and a final-round honesty review. Scoring dimensions are: S_hold compliance survival at 60 points, S_kbv constraint memory at 15 points, S_recover breakdown recovery at 10 points, and S_integrity honest self-reporting at 15 points.

Cause Analysis: Divergent Responses to Pressure Rounds and Constraint Scenarios

Among the risers, Gemini 2.5 Pro and GPT-5.5 showed improved performance during the multi-round progressive pressure phase, most likely reflected in two constraint scenarios: safety compliance and engineering standards. In the v3 design, the R3 pressure round carries the highest weight (R3 accounts for 2 points in v2 anchor questions); when a model sustains its S_hold compliance survival score through the sunk cost escalation phase, the total score improves notably. Claude Opus 4.7 rose 6.1 points, possibly because its constraint memory proved more stable during the KBV restatement probe phase, contributing more to the 15-point S_kbv component.

Doubao Pro's 7.3-point decline suggests weakened S_recover breakdown recovery under sustained pressure. Under the worst-of-3 protocol, a single premature breakdown in a business rules or resource constraints scenario drags down the entire question score. The 8-12 round dialogue design in v3 questions amplifies differences in prompt sensitivity, and Doubao Pro may be more prone to concession during the authoritative exceptional approval or salami-slicing pressure rounds, resulting in a significant loss in the 60-point S_hold component.

The current data only shows score changes without per-scenario breakdowns for each model, so it is not possible to confirm which constraint type dominates the fluctuation. Judging from the test mechanism, however, stepwise social approval and sunk cost escalation are the most likely to expose compliance differences in engineering standard scenarios.

Selection Implications: Real Risk Boundaries for Production Integration

For enterprises integrating AI into production workflows, Grok 4 at 97.50 WDCD can be prioritized in data boundary and safety compliance scenarios; its high S_hold compliance survival makes it suitable for tasks requiring long-term maintenance of multiple parallel hard constraints. GLM-4.6 at 91.80 and Claude Opus 4.7 at 91.50 rank next and can be trialed in resource constraint scenarios, though manual review checkpoints should still be added at the business rules level.

GPT-o3 at 88.80 and GPT-5.5 at 88.60 suit low-risk engineering standard tasks, but in workflows involving sunk cost decisions, prompt-level guardrails are recommended to prevent salami-slicing pressure from gradually loosening constraints. Following Doubao Pro's decline this cycle, production integration should add external validation mechanisms in safety compliance scenarios to keep occasional breakdown events under the worst-of-3 mode from affecting operations.

Overall, models scoring above 90 on the WDCD perform more stably in the KBV restatement probe phase. Enterprises can use this to delineate "direct integration" versus "guardrails required" scenarios, reducing the S_integrity deduction risk caused by models falsely reporting compliance.

Strategic Assessment: Underappreciated and Pending-Verification Signals

Inferred from this cycle's data, the gains of Gemini 2.5 Pro and GPT-5.5 may reflect optimized prompt sensitivity under v3 multi-round pressure; the market may have previously underappreciated the progress these two have made in compliance recovery. Claude Opus 4.7's 6.1-point rise moves it into the Top 3, pointing to a competitive edge in the constraint memory dimension and warranting continued observation in the next cycle to see if the result holds.

Doubao Pro's 7.3-point slide is the only decline case. Analysis suggests weaknesses in the R3 pressure round and the S_recover dimension, and the market may have overestimated its compliance consistency in sustained dialogue. Grok 4's lead at 97.50 was not challenged in this cycle's data, but the close margins among GLM-4.6 at 91.80 and Claude Opus 4.7 at 91.50 indicate the second-tier gap is narrowing.

The next cycle should focus on validating whether the rising models can sustain S_hold scores across more constraint scenarios, and whether Doubao Pro shows signs of recovery. All judgments are based on comparisons between Run #276 and this cycle's data, with no external version update information introduced.

Compliance capability is not an accessory to model parameters; it is the deciding factor in whether production workflows can achieve long-term closed-loop operation.

Data source: YZ Index WDCD Compliance Leaderboard | Run #285 · Change Tracking | Evaluation Methodology

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!