Skip to main content
YZ Index

Evaluation Data

Main Leaderboard WDCD Compliance Test
Currently showing:Run #316 WDCD | 2026-09-09 | Formula v7 | Judge set v6.4
Data Disclosure:To prevent benchmark contamination and overfitting, question texts and expected answers are not disclosed. This page shows model responses, scores, and judging methods for transparency. For the full methodology, seeMethodology page
Model DCD Overall R1 Constraint Acknowledgment R2 Distraction Resistance R3 Constraint Integrity Per Task
Grok 4 grok 93.80 100 100 150
GPT-o3 gpt 89.80 100 100 50
Claude Sonnet 4.6 claude 87.20 100 0 200
DeepSeek V4 Pro deepseek 83.60 100 0 100
Doubao Pro doubao 83.00 100 50 200
GPT-5.5 gpt 80.20 100 50 50
Gemini 3.1 Pro gemini 79.30 100 50 50
Claude Opus 4.7 claude 77.70 100 0 100
Qwen3 Max qwen 71.60 100 100 0
Gemini 2.5 Pro gemini 69.40 100 0 100
GLM-4.6 zhipu 68.00 0 0 0
API Access:For programmatic access to evaluation data, please use our API