CoT Monitoring Faces Plan Injection Attacks; Stanford Paper Shows 25-33% Bypass Rate
A Stanford paper reports that injecting seemingly benign but harmful reasoning plans into an Actor model's context can completely bypass Monitor safety detection in 25-33% of cases. The finding challenges safety strategies that rely on CoT monitoring and suggests that giving the Monitor more compute can sometimes lower detection rates.