Claude Sonnet 4.6's WDCD score fell 5 points from Run #306, while Doubao Pro rose 7.5 points. The opposing movements mark the only significant change in this cycle.
Data Facts: Only Two Models Shifted; the Rest Held Steady
This WDCD v3.1 round evaluated 11 models against a question pool of 29 items. Claude Sonnet 4.6 dropped 5 points from the prior cycle, while Doubao Pro gained 7.5 points. No other model moved by more than 2 points. The current Top 5 is Grok 4 (93.21), GLM-4.6 (91.38), DeepSeek V4 Pro (87.97), GPT-o3 (86.93), and Gemini 3.1 Pro (85.72). The overall adherence score is an equal-weighted average of the v3 native 100-point scale and the v2 anchor-question conversion, sampled under a worst-of-3 rule.
Cause Analysis: Divergent Constraint Memory Under Sustained v3 Pressure
Claude Sonnet 4.6's decline most likely originates in the sustained-pressure phase of the v3 questions. The v3 design spans 8–12 dialogue rounds: it first establishes 2–5 parallel hard constraints, then sequentially applies social-approval, authority-override, salami-slicing, and sunk-cost pressure, and finally scores via a KBV restatement probe and a closing round of honest review. S_hold carries a 60-point weight—the later the model breaks, the higher the score. S_kbv constraint memory is worth 15 points. Claude Sonnet 4.6 appears to have buckled earlier during the R3 pressure round, driving a combined S_hold and S_recover loss exceeding 5 points.
Doubao Pro's 7.5-point gain points to improvements in S_kbv and S_integrity. The v2 anchor questions use a three-round design—R1 constraint injection, R2 interference, R3 pressure—and Doubao Pro appears to have held its constraint memory more consistently through R3, with no false claims of clean compliance. A model that breaks S_integrity yet falsely reports compliance receives 0 points; this cycle, Doubao Pro likely performed better on the KBV restatement probe, lifting its total score directly.
Among the five constraint scenarios, data-boundary and safety-compliance scenarios weighed more heavily on Claude Sonnet 4.6, while resource-limit and engineering-standard scenarios were comparatively friendlier to Doubao Pro. Under worst-of-3 sampling, Claude Sonnet 4.6's worst run saw its breaking point arrive earlier; Doubao Pro's worst run saw it pushed later. This is the direct mechanism behind the opposing score movements.
Selection Implications: Scenario-Tiered Deployment When Integrating AI into Production Workflows
Enterprises integrating AI into production workflows can build a tiered usage strategy from WDCD scores. Grok 4 at 93.21 and GLM-4.6 at 91.38 can be integrated directly in data-boundary and safety-compliance scenarios, given their stronger S_hold resilience and higher post-break recovery probability. DeepSeek V4 Pro at 87.97 suits resource-limit and engineering-standard scenarios, though it warrants additional manual review checkpoints in business-rule scenarios.
With Claude Sonnet 4.6's decline, a second-confirmation mechanism is advisable in multi-round incremental-pressure settings, particularly when conversations exceed eight rounds. With Doubao Pro's gain, its use can be extended across workflows with heavier R3 pressure rounds, while retaining S_recover recovery-path monitoring. GPT-o3 and Gemini 3.1 Pro remain stable and can serve as neutral fallbacks, invoked in parallel with the top two models in mixed scenarios.
Strategic Assessment: Signals of Underestimated and Overestimated Adherence Capability
Doubao Pro's 7.5-point gain this cycle may be underestimated by the market. Its steady performance through the R3 pressure phase of the v2 anchor questions suggests its constraint-memory capability is approaching GLM-4.6's level. If it maintains or slightly extends this next cycle, Doubao Pro could enter the Top 4. Claude Sonnet 4.6's 5-point drop, by contrast, signals vulnerability under the combined social-approval and sunk-cost pressure—a finding worth prioritizing for verification next cycle.
Grok 4 and GLM-4.6's leading positions went unchallenged this cycle, indicating they still hold the edge on both the KBV restatement probe and the closing honest self-report. DeepSeek V4 Pro at 87.97 trails GPT-o3 at 86.93 by just 1.04 points; if DeepSeek V4 Pro gains another 1–2 points on S_integrity next cycle, it could overtake GPT-o3.
The analysis suggests that Claude Sonnet 4.6 is the model whose adherence capability the market overestimates, while Doubao Pro is the one it underestimates. The signals worth validating next cycle are whether Doubao Pro can sustain its R3 stability through the salami-slicing phase of v3 questions, and whether Claude Sonnet 4.6 will slide further in data-boundary scenarios.
Adherence scores are not static labels but dynamic survival curves under multi-round pressure; the true watershed next cycle will emerge after round eight.
Data source: YZ Index WDCD Adherence Leaderboard | Run #311 · Change Tracking | Evaluation Methodology
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接