WDCD Run #161: Average Instruction Decay Hits -48.6% Across 11 Models, GPT-5.5 Leads at 89.2 Points
WDCD Run #161 (2026-06-11) evaluated 11 large language models on multi-turn commitment integrity, recording an average instruction decay of -48.6% from Round 1 to Round 3. GPT-5.5 led the ranking with 89.2 points, while Doubao Pro showed the strongest decay resistance.