DCD · Dynamic Context Decay
After 5000 characters of distraction, does the AI still remember what you said three minutes ago?
DCD Leaderboard
| # | Model | WDCD | R1 Understanding | R2 Resistance | R3 Integrity | Main Score | vs Main Rank |
|---|---|---|---|---|---|---|---|
| #1 | DeepSeek V4 Pro | 97.7 | 100% | 100% | 100% | 69.9 | ↑8 |
| #2 | Gemini 3.1 Pro | 96.7 | 100% | 100% | 100% | 70.7 | ↑5 |
| #3 | Grok 4 | 96.0 | 100% | 100% | 50% | 79.6 | ↓1 |
| #4 | Gemini 2.5 Pro | 94.5 | 100% | 0% | 100% | 65.3 | ↑6 |
| #5 | GPT-o3 | 94.2 | 100% | 0% | 50% | 76.6 | — |
| #6 | Claude Opus 4.7 | 92.6 | 100% | 0% | 100% | 80.1 | ↓5 |
| #7 | GPT-5.5 | 90.1 | 100% | 0% | 100% | 77.8 | ↓4 |
| #8 | Claude Sonnet 4.6 | 83.6 | 100% | 0% | 100% | 76.8 | ↓4 |
| #9 | Doubao Pro | 76.7 | 100% | 100% | 100% | 72.0 | ↓3 |
| #10 | GLM-4.6 | 76.5 | 0% | 0% | 0% | — | — |
| #11 | Qwen3 Max | 75.2 | 100% | 0% | 100% | 70.0 | ↓3 |
Three-Round Constraint Retention Curve (v2 anchor questions)
Each row represents a model. The three color bars represent the scoring rates for R1 (understanding), R2 (anti-interference), R3 (constraint adherence) respectively.
DeepSeek V4 Pro
100%
Gemini 3.1 Pro
100%
Grok 4
50%
Gemini 2.5 Pro
100%
GPT-o3
50%
Claude Opus 4.7
100%
GPT-5.5
100%
Claude Sonnet 4.6
100%
Doubao Pro
100%
GLM-4.6
0%
Qwen3 Max
100%
R1 R1 Understanding
R2 R2 Resistance
R3 R3 Integrity
Performance Across Five Constraint Types
Which model is most likely to fail under which type of constraint?
| Model | Data Boundary | Resource Limit | Business Rule | Security | Engineering |
|---|---|---|---|---|---|
| DeepSeek V4 Pro | 100 | 100 | 100 | 100 | 89 |
| Gemini 3.1 Pro | 100 | 88 | 100 | 96 | 100 |
| Grok 4 | 88 | 93 | 100 | 100 | 100 |
| Gemini 2.5 Pro | 88 | 93 | 100 | 93 | 100 |
| GPT-o3 | 75 | 100 | 96 | 100 | 100 |
| Claude Opus 4.7 | 88 | 100 | 100 | 100 | 75 |
| GPT-5.5 | 88 | 100 | 81 | 100 | 81 |
| Claude Sonnet 4.6 | 83 | 81 | 70 | 100 | 84 |
| Doubao Pro | 98 | 70 | 56 | 75 | 84 |
| GLM-4.6 | 43 | 85 | 85 | 85 | 85 |
| Qwen3 Max | 88 | 70 | 75 | 56 | 86 |
🏛 Why We Made WDCD
Design philosophy, differences from existing evaluations, roadmap — the complete story of the world's first multi-round commitment evaluation framework.
📋 Methodology
How does WDCD test? How is multi-turn escalating pressure designed? How is scoring 100% auditable?
📊 API Interface
Get WDCD raw data for third-party research and visualization.
📰 All Cases
Complete list of all cases where R1 passed but R3 collapsed.