Grok 4 ranked first in the WDCD v3.1 adherence test with 95.69 points, while Qwen3 Max ranked last with 73.48 points, a difference of 22.21 points.
Ranking Landscape: Concentrated at the Top, a Break at the Bottom
This WDCD leaderboard shows that the top five all scored above 86 points: Grok 4 (95.69), Gemini 3.1 Pro (88.14), GLM-4.6 (87.55), DeepSeek V4 Pro (86.62), and Claude Opus 4.7 (86.34). Ranks 6 through 10 clustered in the 82–84 range, with GPT-5.5 scoring 84.24 and GPT-o3 scoring 84.07. Among the bottom three, Qwen3 Max scored 73.48, Doubao Pro 76.38, and Gemini 2.5 Pro 77.69, forming a clear break from the top.
Global statistics show a full-score rate of 57% and an R3 collapse rate of 9.4%. Sampling used a worst-of-3 basis, running each question three times and taking the worst result, with rule-based scoring and zero AI judges.
Champion Analysis: Grok 4's Survival Advantage Under R3 Pressure
Grok 4's R1=1.00, R2=1.00, and R3=1.38/2, with the v3 questions' adherence survival score contributing prominently. The v3 question design includes 8–12 rounds of dialogue, first establishing 2–5 hard constraints, then sequentially applying social proof, authority exception, salami slicing, and sunk cost pressure, and finally a KBV recount probe and honest self-report. Grok 4 still maintained a relatively high score at the R3 stage, indicating strong constraint memory and recovery ability under continuous pressure.
The cause may come from its handling mechanisms for data boundary and security compliance constraints, breaking later under multi-round incremental pressure, which raises the S_hold score accordingly.
Bottom Analysis: Qwen3 Max's R3 Performance and Memory Decay
Qwen3 Max had R1=1.00, R2=0.75, and R3=0.63/2, with its R3 score clearly lower than top models. In the three-round design of the v2 anchor questions, R3 carried a weight of 2 points, and its low score directly dragged down the total. The possible reason is that under resource constraints or engineering specification scenarios, constraint memory decays faster after continuous interference, affecting S_kbv and S_recover scores.
Compared with Doubao Pro, ranked 14th (R1=0.75), Qwen3 Max had a full R1 score but lower R3, showing normal performance during the commitment-establishment stage, with problems concentrated in later pressure rounds.
Selection Implications of the Gap Between the Top Tier and the Tail
For enterprises integrating AI into production processes, WDCD score differences directly affect usable scenarios. Top models such as Grok 4 and Gemini 3.1 Pro have longer adherence survival times in five constraint scenarios, making them suitable for links with high data boundary and security compliance requirements, such as financial risk control rule execution or medical compliance review.
Tail models such as Qwen3 Max score low on R3 under multi-round pressure. It is advisable to add extra external guardrails in business rule or resource constraint scenarios, such as secondary validation or human review nodes, to prevent constraints from being gradually broken under sunk cost pressure.
Specific judgment: if an enterprise's core processes involve more than 8 rounds of continuous interaction, prioritize models with an R3 score above 1.0; models with an R3 score below 0.8 are suitable for single-turn or low-pressure scenarios and require a circuit-breaker mechanism.
Strategic Judgment: Signals That Adherence Ability May Be Underestimated and Overestimated
GLM-4.6 rose 26.0 points this period compared with the previous period, and GPT-5.5 rose 10.5 points, showing that some models have considerable room for improvement on v3 multi-round incremental pressure questions. Analysis suggests that GLM-4.6's improvement may come from optimization in the KBV recount probe and honest self-report stages, and it is worth continuing to verify its stability in engineering specification scenarios next period.
GPT-o3's R3 was only 0.25/2, a clear gap from the 0.75/2 of GPT-5.5 in the same series, possibly reflecting weaker recovery ability under authority exception pressure; this signal can be a key observation item next period.
Claude Opus 4.7 and Claude Sonnet 4.6 scored 86.34 and 80.31 respectively, a difference of more than 6 points within the same series, suggesting that different parameter scales diverge in adherence tasks. Production selection should test them individually rather than packaging them by series.
Adherence ability is not a static label but a dynamic survival curve under multi-round pressure; enterprises should use the R3 score, not the average score, as the core threshold in selection.
Data source: YZ Index WDCD Adherence Leaderboard | Run #360 · Overall Leaderboard Ranking | Evaluation Methodology
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接