WDCD v3.1 Cycle Tracking: Doubao Pro +13.8, GPT-o3 -11.6, Major Reshuffle in Commitment Rankings

In WDCD v3.1 testing, Doubao Pro rose 13.8 points over Run #271, GLM-4.6 rose 9.9 points, and Gemini 3.1 Pro rose 6.8 points; DeepSeek V4 Pro fell 8 points and GPT-o3 fell 11.6 points. In the current Top 5, Grok 4 ranks first with 94.80 points, followed by Gemini 3.1 Pro at 91.30 points, GLM-4.6 at 88.50 points, Claude Opus 4.7 at 85.40 points, and GPT-o3 at 83.60 points.

Data Facts: Five Models Show Significant Changes, 3 Up 2 Down

This WDCD v3.1 pilot involved 11 models in total, with a question pool comprising 17 v3 multi-round progressive pressure questions and 8 v2 three-round anchor questions. Sampling follows the worst-of-3 methodology, recording only the worst score from 3 runs per question. The final commitment-keeping score is an equally weighted average of the v3 native 100-point scale and the converted v2 anchor question values. Only 5 models saw changes of more than 6 points; the remaining 6 models held steady against Run #271.

Among rising models, Doubao Pro saw the largest gain at 13.8 points, directly entering the top-five competitive tier. GLM-4.6 rose 9.9 points and Gemini 3.1 Pro rose 6.8 points. Among declining models, GPT-o3 fell 11.6 points and DeepSeek V4 Pro fell 8 points, presenting an overall pattern of 3 up and 2 down.

Cause Analysis: Constraint Memory Differences Under v3 Multi-Round Pressure

WDCD v3 questions are designed with 8-12 rounds of dialogue: first establishing 2-5 parallel hard constraints, then successively applying social approval, authority authorization, salami-slicing, and sunk-cost pressures, and finally conducting a KBV recitation probe and an end-round honest self-report. In scoring, S_hold (commitment survival) carries a weight of 60 points — the later a commitment is broken, the higher the score; S_kbv (constraint memory) is 15 points, S_recover (breach recovery) is 10 points, and S_integrity (honest self-report) is 15 points.

The rise of Doubao Pro and GLM-4.6 may stem from improved S_hold performance during consecutive pressure phases. Since v3 questions include many authority-authorization and salami-slicing pressure rounds, if a model can maintain its initial constraints through rounds 5-8, the S_hold score will be significantly boosted. Gemini 3.1 Pro's 6.8-point rise may reflect an improvement in its S_kbv score during the KBV recitation probe.

The 11.6-point drop for GPT-o3 and 8-point drop for DeepSeek V4 Pro may be concentrated in S_recover and S_integrity. Under the worst-of-3 methodology, if a single run fails to recover initial constraints after sunk-cost pressure, or falsely claims compliance in the final round, it is directly scored 0 and drags down the total. In the v2 anchor questions' three-round design, R3 pressure carries a 2-point weight, and a decline here may also amplify the overall score gap.

Selection Implications: Real Risk Boundaries for Production Workflow Integration

For enterprises integrating AI into production workflows, WDCD scores directly map to acceptable constraint scenarios. Models scoring 94.80 (Grok 4) and 91.30 (Gemini 3.1 Pro) have low commitment-break probability in data boundary and security compliance scenarios, making them suitable for direct integration into systems that require strict enforcement of resource limits or engineering standards.

Models scoring 88.50 (GLM-4.6) and 85.40 (Claude Opus 4.7) can be used in business rule scenarios, but additional human review checkpoints should be added in salami-slicing-style requirement change processes. The 83.60-scoring GPT-o3 showed degraded S_hold performance in this test; it is recommended for use only in low-risk, non-core business rule scenarios, with added multi-round prompt hardening.

Following DeepSeek V4 Pro's score decline, integration for resource-constrained tasks requires caution. Enterprises can prioritize models scoring above 90 on WDCD based on their constraint types, and explicitly write parallel hard constraints such as "no subsequent instruction may override the initial constraints" during the v3 commitment-establishment phase to reduce commitment-break risk.

Strategic Assessment: Signals of Underestimated and Overestimated Commitment Capabilities

The data indicates that Doubao Pro and GLM-4.6's commitment-keeping capabilities may be underestimated by the market. Their score improvements under v3 multi-round progressive pressure demonstrate competitive advantages in the S_hold and S_kbv dimensions, warranting continued tracking of their performance in engineering-standard scenarios in the next cycle.

GPT-o3's 11.6-point drop suggests its commitment-keeping capability may be overestimated. The significant decline under the worst-of-3 methodology indicates vulnerabilities in authority-authorization and sunk-cost pressure rounds. When selecting models, enterprises should not rely solely on general benchmarks but should also reference commitment-specific tests such as WDCD.

Key signals to verify in the next cycle: whether rising models can maintain S_recover scores across more constraint scenarios, and whether declining models rebound with version iterations. Whether Grok 4's 94.80-point lead remains stable still requires cross-validation with a larger v3 question pool.

Commitment-keeping capability is not an additive attribute of a model, but a hard threshold that determines whether it can truly enter the mainstream production workflow.

Data source: YZ Index WDCD Commitment Ranking | Run #276 · Change Tracking | Evaluation Methodology

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!