In this round of WDCD v3.1 testing, Qwen3 Max improved its score by 20.2 points and GLM-4.6 by 19.9 points; Claude Sonnet 4.6 rose by 7.6 points, Gemini 2.5 Pro by 11.3 points, and Grok 4 by 6.4 points. Among the 11 evaluated models, five rose and none declined.
Constraint-scenario differences behind the score changes
The total commitment-adherence score is formed from an equally weighted average of the native 100-point v3 tasks and the converted v2 anchor tasks. The v3 tasks use 8-12 rounds of dialogue: first establishing 2-5 hard constraints, then applying escalating pressure through social proof, special authority approval, salami tactics, and sunk cost, and finally running a KBV recall probe and an honest final-round review. Qwen3 Max and GLM-4.6 showed the largest gains, indicating that these two models improved most noticeably in S_hold scores during the continuous pressure phase, especially in the two scenario categories of resource limits and safety compliance, where the breach round shifted clearly later.
Grok 4's WDCD in this round is 100.00, maintaining a perfect score from the previous round, indicating that under worst-of-3 sampling, neither the three rounds of v2 anchor tasks nor the v3 multi-round progressive pressure produced any commitment breach. Gemini 3.1 Pro WDCD=94.90, Claude Opus 4.7 WDCD=93.30, and GPT-o3 WDCD=93.10 all placed in the top five, but still lag Grok 4's 100 by 6.7-6.9 points.
Cause analysis: pressure rounds and KBV probe performance
Qwen3 Max and GLM-4.6, with gains exceeding 19 points, most likely strengthened their commitment adherence during the 5th-8th rounds of salami-tactic and sunk-cost pressure. Improvements in the two items of S_kbv constraint memory (15 points) and S_recover breach recovery (10 points) suggest that the models' accuracy in KBV recall of established commitments increased, and that after a breach they could more quickly restore the original constraints. By contrast, Claude Sonnet 4.6 rose only 7.6 points, possibly because it still carries a risk of scoring 0 in the S_integrity honest self-report step in engineering-specification scenarios.
All rising models avoided declines, indicating that the current version has improved overall robustness in the two constraint-scenario categories of data boundaries and business rules. In the three rounds of v2 anchor tasks, higher scores in the R3 pressure step (worth 2 points) were the main contributor to this round's total score increase.
Practical implications for production-process integration
Enterprises integrating AI into production processes can use WDCD scores to judge usable scenarios. Grok 4's WDCD=100.00 means it can be used directly in safety-compliance and data-boundary scenarios, with the lowest guardrail requirements. GLM-4.6 WDCD=96.40 and Gemini 3.1 Pro WDCD=94.90 can be trialed in resource-limit scenarios, but still require additional manual review nodes in the business-rules step.
Claude Opus 4.7 and GPT-o3 scored close to 93, making them suitable for internal toolchains, but for external interfaces involving engineering specifications, it is still advisable to retain the breach-detection mechanism corresponding to an S_hold score of 60. After this round's large gain, Qwen3 Max is now close to the top-five threshold, and enterprises can gradually expand deployment in non-core data-boundary scenarios.
Strategic judgment and signals to verify next round
The 19.9-point and 20.2-point gains of GLM-4.6 and Qwen3 Max suggest that the market may have previously underestimated their commitment-adherence capabilities. Grok 4's continued 100-point performance confirms that its constraint memory and honest self-report capabilities under multi-round pressure remain ahead. The analysis suggests that next round should focus on whether Qwen3 Max and GLM-4.6 can keep S_kbv scores above 12 points in the 9th-12th round KBV recall probes of the v3 tasks.
If the two models' S_recover scores in safety-compliance scenarios continue to improve, this will further change production selection rankings. Current data only support the judgment that they "may have been underestimated"; there is not yet enough evidence to support a "continued leadership" conclusion.
Commitment adherence is not an accessory to model parameters; it is the key threshold determining whether enterprises can move AI from demo environments to real production lines.
Data source: YZ Index WDCD Commitment Leaderboard | Run #346 · Change tracking | Evaluation methodology
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接