DCD · Dynamic Context Decay
After 5000 characters of distraction, does the AI still remember what you said three minutes ago?
DCD Leaderboard
| # | Model | WDCD | R1 Understanding | R2 Resistance | R3 Integrity | Main Score | vs Main Rank |
|---|---|---|---|---|---|---|---|
| #1 | Grok 4 | 94.8 | 100% | 100% | 75% | 76.4 | ↑3 |
| #2 | DeepSeek V4 Pro | 93.6 | 100% | 100% | 75% | 79.3 | — |
| #3 | GLM-4.6 | 93.5 | 100% | 100% | 50% | 53.1 | ↑7 |
| #4 | Claude Opus 4.7 | 92.6 | 100% | 100% | 75% | 82.6 | ↓3 |
| #5 | Claude Sonnet 4.6 | 88.2 | 100% | 50% | 50% | 72.9 | ↑1 |
| #6 | GPT-o3 | 85.7 | 100% | 100% | 25% | 76.7 | ↓3 |
| #7 | Gemini 3.1 Pro | 81.0 | 100% | 100% | 50% | 70.8 | ↑1 |
| #8 | GPT-5.5 | 73.9 | 100% | 50% | 0% | 75.9 | ↓3 |
| #9 | Gemini 2.5 Pro | 67.4 | 100% | 50% | 75% | 70.0 | — |
| #10 | Qwen3 Max | 66.7 | 100% | 100% | 0% | 71.5 | ↓3 |
| #11 | Doubao Pro | 64.2 | 50% | 100% | 25% | 52.2 | — |
Three-Round Constraint Retention Curve (v2 anchor questions)
Each row represents a model. The three color bars represent the scoring rates for R1 (understanding), R2 (anti-interference), R3 (constraint adherence) respectively.
Performance Across Five Constraint Types
Which model is most likely to fail under which type of constraint?
| Model | Data Boundary | Resource Limit | Business Rule | Security | Engineering |
|---|---|---|---|---|---|
| Grok 4 | 100 | 100 | 88 | 100 | 86 |
| DeepSeek V4 Pro | 80 | 100 | 88 | 100 | 100 |
| GLM-4.6 | 100 | 93 | 75 | 100 | 100 |
| Claude Opus 4.7 | 100 | 100 | 88 | 100 | 75 |
| Claude Sonnet 4.6 | 95 | 100 | 75 | 88 | 84 |
| GPT-o3 | 100 | 100 | 88 | 75 | 66 |
| Gemini 3.1 Pro | 100 | 69 | 71 | 96 | 69 |
| GPT-5.5 | 81 | 81 | 56 | 63 | 88 |
| Gemini 2.5 Pro | 100 | 80 | 63 | 69 | 25 |
| Qwen3 Max | 81 | 70 | 75 | 45 | 61 |
| Doubao Pro | 98 | 65 | 38 | 40 | 81 |
Notable Failure Cases
R1 confirms understanding of the constraint → R3 fully compromises cases (dialogue desensitized display).
🏛 Why We Made WDCD
Design philosophy, differences from existing evaluations, roadmap — the complete story of the world's first multi-round commitment evaluation framework.
📋 Methodology
How does WDCD test? How is multi-turn escalating pressure designed? How is scoring 100% auditable?
📊 API Interface
Get WDCD raw data for third-party research and visualization.
📰 All Cases
Complete list of all cases where R1 passed but R3 collapsed.