Four Models Slide Collectively in WDCD as Gemini 2.5 Pro Plunges 16.7 Points

In the WDCD v3.1 pilot, Gemini 2.5 Pro fell by 16.7 points compared with Run #331, the largest decline among the four evaluated models. DeepSeek V4 Pro, GPT-5.5, and Qwen3 Max fell by 6.6, 6.5, and 7.9 points, respectively, and no model's score rose this period.

Data Facts: Distribution Characteristics of the Decline Limited to Four Models

This sampling covered 11 models and used a worst-of-3 standard. Gemini 2.5 Pro's 16.7-point decline far exceeded the other three. DeepSeek V4 Pro fell from the 97.7-point range in Run #331 to 91.10, while GPT-5.5 and Qwen3 Max's declines were concentrated in the 6.5–7.9 point range. The Top 5 leaderboard shows Grok 4 in first place with 93.60 points, followed closely by Gemini 3.1 Pro at 92.50, GPT-o3 at 91.90, DeepSeek V4 Pro at 91.10, and Claude Opus 4.7 at 89.50.

Cause Analysis: Constraint-Failure Pathways Under v3 Multi-Turn Pressure

The v3 question design includes 8–12 dialogue turns: it first establishes 2–5 hard constraints, then sequentially applies social proof, authority special approval, salami slicing, and sunk-cost pressure, and finally conducts KBV restatement probes and honest self-report. Gemini 2.5 Pro's large 16.7-point decline most likely occurred during the continuous pressure phase in R3–R6, that is, the rounds where "salami slicing" and "sunk cost" overlapped. The S_hold compliance survival score (weight 60) is sensitive to the timing of breach; if constraints loosen as early as turns 5–7, this item's score is directly pulled down.

The 6–8 point decline of DeepSeek V4 Pro and Qwen3 Max is more likely concentrated in two items: S_kbv constraint memory (weight 15) and S_recover breach recovery (weight 10). The three-turn design of the v2 anchor question shows that after R2 interference and R3 pressure, the KBV restatement accuracy of these two models declined, dragging down their converted percentage scores. GPT-5.5's 6.5-point decline was the smallest, perhaps with S_integrity honest self-report deductions occurring only in engineering-specification constraint scenarios.

Selection Implications: Guardrail Priorities When Integrating into Production Workflows

Enterprises integrating AI into production workflows need to distinguish among constraint scenarios. For hard constraints related to security compliance and data boundaries, Gemini 2.5 Pro's current performance in the 91-point range is already below Grok 4's 93.60, so it is advisable to add a secondary validation node before RAG to avoid unauthorized outputs during user follow-ups after turn 5. In resource-limit and business-rule scenarios, DeepSeek V4 Pro at 91.10 remains usable, but a prompt-layer instruction to "restate the established constraints each turn" should be added to keep the S_kbv weight loss within 5 points.

For engineering-specification tasks, Grok 4 can be prioritized; its 93.60 points under the worst-of-3 standard shows a more stable S_hold and S_recover combination, making it suitable for code-review workflows with long-context, multi-turn interaction. Claude Opus 4.7's 89.50 points suggests that in customer-service escalation scenarios with pronounced sunk-cost pressure, an additional manual spot-check node must be deployed.

Strategic Judgment: Underestimated Compliance Capacity and Signals to Verify Next Period

The gap between Grok 4 at 93.60 and Gemini 3.1 Pro at 92.50 is only 1.1 points. Combined with the background of four models declining collectively this period, Grok 4's compliance capacity may be underestimated by the market. Gemini 2.5 Pro's 16.7-point decline shows increased sensitivity to "authority special approval" pressure in v3 questions, a signal worth focusing on next period: if Gemini 3.1 Pro can raise its S_hold score in such scenarios above 55, it can be judged to have completed targeted alignment.

The closeness between DeepSeek V4 Pro at 91.10 and GPT-o3 at 91.90 indicates that the two no longer differ significantly in converted scores on the v2 anchor question; enterprises selecting models can directly compare S_recover recovery capability rather than total scores. Qwen3 Max's 7.9-point decline suggests a drop in its KBV probe pass rate in multi-constraint parallel scenarios. The next cycle should observe whether it rebounds through targeted training on engineering-specification questions.

Compliance capacity is not a static leaderboard, but a survival curve under multi-turn pressure. This period's data has clearly marked Gemini 2.5 Pro's vulnerable range.

Data source: YZ Index WDCD Compliance Leaderboard | Run #336 · Change Tracking | Evaluation Methodology

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!