Skip to main content
🧪

Experimental Dimension

DCD (Dynamic Context Decay) tests whether AI maintains constraints across multi-turn conversations. Data is still being collected and does not count towards the main leaderboard.

Learn about the methodology →

DCD · Dynamic Context Decay

After 5000 characters of distraction, does the AI still remember what you said three minutes ago?

Winzheng Index v7 experimental dimension · 25 questions · multi-turn escalating pressure · 11 models

DCD Leaderboard

# Model WDCD R1 Understanding R2 Resistance R3 Integrity Main Score vs Main Rank
#1 DeepSeek V4 Pro 97.7 100% 100% 100% 69.9 ↑8
#2 Gemini 3.1 Pro 96.7 100% 100% 100% 70.7 ↑5
#3 Grok 4 96.0 100% 100% 50% 79.6 ↓1
#4 Gemini 2.5 Pro 94.5 100% 0% 100% 65.3 ↑6
#5 GPT-o3 94.2 100% 0% 50% 76.6
#6 Claude Opus 4.7 92.6 100% 0% 100% 80.1 ↓5
#7 GPT-5.5 90.1 100% 0% 100% 77.8 ↓4
#8 Claude Sonnet 4.6 83.6 100% 0% 100% 76.8 ↓4
#9 Doubao Pro 76.7 100% 100% 100% 72.0 ↓3
#10 GLM-4.6 76.5 0% 0% 0%
#11 Qwen3 Max 75.2 100% 0% 100% 70.0 ↓3

Three-Round Constraint Retention Curve (v2 anchor questions)

Each row represents a model. The three color bars represent the scoring rates for R1 (understanding), R2 (anti-interference), R3 (constraint adherence) respectively.

DeepSeek V4 Pro
100%
Gemini 3.1 Pro
100%
Grok 4
50%
Gemini 2.5 Pro
100%
GPT-o3
50%
Claude Opus 4.7
100%
GPT-5.5
100%
Claude Sonnet 4.6
100%
Doubao Pro
100%
GLM-4.6
0%
Qwen3 Max
100%
R1 R1 Understanding R2 R2 Resistance R3 R3 Integrity

Performance Across Five Constraint Types

Which model is most likely to fail under which type of constraint?

Model Data Boundary Resource Limit Business Rule Security Engineering
DeepSeek V4 Pro 100 100 100 100 89
Gemini 3.1 Pro 100 88 100 96 100
Grok 4 88 93 100 100 100
Gemini 2.5 Pro 88 93 100 93 100
GPT-o3 75 100 96 100 100
Claude Opus 4.7 88 100 100 100 75
GPT-5.5 88 100 81 100 81
Claude Sonnet 4.6 83 81 70 100 84
Doubao Pro 98 70 56 75 84
GLM-4.6 43 85 85 85 85
Qwen3 Max 88 70 75 56 86