YZ Index
Evaluation Data
Currently showing:Run #316 WDCD | 2026-09-09 | Formula v7 | Judge set v6.4
Data Disclosure:To prevent benchmark contamination and overfitting, question texts and expected answers are not disclosed. This page shows model responses, scores, and judging methods for transparency. For the full methodology, seeMethodology page。
| Model | DCD Overall | R1 Constraint Acknowledgment | R2 Distraction Resistance | R3 Constraint Integrity | Per Task |
|---|---|---|---|---|---|
| Grok 4 grok | 93.80 | 100 | 100 | 150 | |
| GPT-o3 gpt | 89.80 | 100 | 100 | 50 | |
| Claude Sonnet 4.6 claude | 87.20 | 100 | 0 | 200 | |
| DeepSeek V4 Pro deepseek | 83.60 | 100 | 0 | 100 | |
| Doubao Pro doubao | 83.00 | 100 | 50 | 200 | |
| GPT-5.5 gpt | 80.20 | 100 | 50 | 50 | |
| Gemini 3.1 Pro gemini | 79.30 | 100 | 50 | 50 | |
| Claude Opus 4.7 claude | 77.70 | 100 | 0 | 100 | |
| Qwen3 Max qwen | 71.60 | 100 | 100 | 0 | |
| Gemini 2.5 Pro gemini | 69.40 | 100 | 0 | 100 | |
| GLM-4.6 zhipu | 68.00 | 0 | 0 | 0 |
API Access:For programmatic access to evaluation data, please use our API。