Skip to main content
YZ Index

Evaluation Data

Main Leaderboard WDCD Compliance Test
Currently showing:Run #253 WDCD | 2026-07-29 | Formula v7 | Judge set v6.4
Data Disclosure:To prevent benchmark contamination and overfitting, question texts and expected answers are not disclosed. This page shows model responses, scores, and judging methods for transparency. For the full methodology, seeMethodology page
Model DCD Overall R1 Constraint Acknowledgment R2 Distraction Resistance R3 Constraint Integrity Per Task
Grok 4 grok 94.80 100 100 150
DeepSeek V4 Pro deepseek 93.60 100 100 150
GLM-4.6 zhipu 93.50 100 100 100
Claude Opus 4.7 claude 92.60 100 100 150
Claude Sonnet 4.6 claude 88.20 100 50 100
GPT-o3 gpt 85.70 100 100 50
Gemini 3.1 Pro gemini 81.00 100 100 100
GPT-5.5 gpt 73.90 100 50 0
Gemini 2.5 Pro gemini 67.40 100 50 150
Qwen3 Max qwen 66.70 100 100 0
Doubao Pro doubao 64.20 50 100 50
API Access:For programmatic access to evaluation data, please use our API