YZ Index
Evaluation Data
Currently showing:Run #253 WDCD | 2026-07-29 | Formula v7 | Judge set v6.4
Data Disclosure:To prevent benchmark contamination and overfitting, question texts and expected answers are not disclosed. This page shows model responses, scores, and judging methods for transparency. For the full methodology, seeMethodology page。
| Model | DCD Overall | R1 Constraint Acknowledgment | R2 Distraction Resistance | R3 Constraint Integrity | Per Task |
|---|---|---|---|---|---|
| Grok 4 grok | 94.80 | 100 | 100 | 150 | |
| DeepSeek V4 Pro deepseek | 93.60 | 100 | 100 | 150 | |
| GLM-4.6 zhipu | 93.50 | 100 | 100 | 100 | |
| Claude Opus 4.7 claude | 92.60 | 100 | 100 | 150 | |
| Claude Sonnet 4.6 claude | 88.20 | 100 | 50 | 100 | |
| GPT-o3 gpt | 85.70 | 100 | 100 | 50 | |
| Gemini 3.1 Pro gemini | 81.00 | 100 | 100 | 100 | |
| GPT-5.5 gpt | 73.90 | 100 | 50 | 0 | |
| Gemini 2.5 Pro gemini | 67.40 | 100 | 50 | 150 | |
| Qwen3 Max qwen | 66.70 | 100 | 100 | 0 | |
| Doubao Pro doubao | 64.20 | 50 | 100 | 50 |
API Access:For programmatic access to evaluation data, please use our API。