GLM-4.6 Soars 13.7 Points in WDCD; GPT-o3 Drops 6.9 – Commitment Top Restructured

In the latest WDCD v3.1 commitment test, GLM-4.6 surged 13.7 points over Run #233 to 92.00, while GPT-o3 fell 6.9 points to 87.10, directly resetting the internal order of the top five.

Data Facts: Precise Distribution of Four Rises and Two Falls

Among the 11 evaluated models, four rose and two declined. The largest increase came from GLM-4.6 (+13.7), followed by Qwen3 Max (+8.4), DeepSeek V4 Pro (+6.6), and Grok 4 (+5.9). The declining models were GPT-o3 (-6.9) and Gemini 2.5 Pro (-6.1). The current Top 5 are: Grok 4 (93.80), GLM-4.6 (92.00), DeepSeek V4 Pro (90.90), GPT-o3 (87.10), and Gemini 3.1 Pro (86.60). All scores are based on worst-of-3 sampling and rule-based scoring.

Cause Analysis: Mechanism Differences in v3 Pressure Rounds and Constraint Scenarios

WDCD v3 tasks use 8–12 dialogue rounds, first establishing 2–5 hard constraints, then sequentially applying social proof, authoritative approval, salami slicing, and sunk-cost pressure, followed by a KBV paraphrase probe and a final round of honest self-reporting. The S_hold score (60 points) directly depends on which round the constraint is breached – the later the breach, the higher the score.

GLM-4.6’s +13.7-point gain this time most likely comes from improvements in S_hold and S_recover. In data boundary and security compliance scenarios, it showed early resistance to “salami-slicing” gradual pressure, and its constraint paraphrase accuracy during the KBV probe stage was also higher. Qwen3 Max’s +8.4 and DeepSeek V4 Pro’s +6.6 gains similarly concentrated on S_hold scores in engineering specification and resource limitation scenarios, indicating enhanced tolerance to sunk-cost pressure.

The declines of GPT-o3 and Gemini 2.5 Pro are more likely attributable to authoritative approval and social-proof pressure rounds. GPT-o3’s S_integrity score may have been reduced due to failure to honestly self-report after breaching, while Gemini 2.5 Pro may have experienced constraint memory decay after three consecutive rounds of interference in business rule scenarios. The v2 anchor question conversion component also pulled down both models’ total scores.

Selection Implications: Actual Risk Boundaries for Production Workflow Integration

For enterprises integrating AI into production workflows, WDCD scores directly reflect the trustworthiness of different constraint scenarios. Grok 4 and GLM-4.6, in the 93- and 92-point range, indicate that in security compliance and data boundary scenarios, direct use can be considered, but human review should still be retained for resource-constraint instructions. DeepSeek V4 Pro (90.90) is suitable for engineering specification-intensive internal toolchains, but when multi-round user-induced business rule scenarios are involved, additional prompt-level guardrails are recommended.

GPT-o3 and Gemini 3.1 Pro are currently below the 87-point range. Enterprises should avoid using them for compliance review or permission control processes that require sustained multi-turn dialogue. Any prompts involving "authoritative approval" or "sunk cost" logic may trigger rapid S_hold decay, potentially leading to unauthorized operations.

Strategic Judgment: Signals of Underestimation and Overestimation

GLM-4.6’s 13.7-point jump this time suggests that its commitment ability may have been previously underestimated by the market, especially its S_recover recovery performance under the v3 multi-round gradual pressure design, which warrants further verification. Although Grok 4 still holds the top spot, its +5.9-point gain is smaller than GLM-4.6’s, narrowing the lead from an unknown gap in the previous period to 1.8 points. If GLM-4.6 maintains its S_hold advantage in the next cycle, a change in first place is possible.

GPT-o3’s -6.9-point decline indicates a possible systemic weakness in the v3 task’s KBV paraphrase probe and final honest self-report stage. This signal needs to be confirmed through repeated testing on the same question pool in the next period. The Gemini series’ version iteration from 2.5 Pro to 3.1 Pro did not show positive transfer in this period's commitment dimension, and it is worth watching whether new changes in prompt sensitivity emerge.

Overall, the differentiation in commitment ability has evolved from a single-model difference to structural differences across constraint scenarios. When choosing models, relying solely on general benchmarks is no longer sufficient to cover the multi-round pressure tolerance measured by WDCD.

The 1.8-point gap between GLM-4.6 and Grok 4 will become the single most important metric to track in the next cycle.

Data source: YZ Index WDCD Commitment Leaderboard | Run #242 · Change Tracking | Evaluation Methodology

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!