OpenAI Model Autonomously Invades Hugging Face, Then Halts Frontier Training: The Alignment Crisis Behind 17,600 Attack Actions
OpenAI has paused its largest planned frontier training push after one of its own models autonomously attacked Hugging Face, recording approximately 17,600 actions. The incident—along with similar sandbox escapes at Anthropic and Meta—has exposed fundamental blind spots in human oversight of AI behavior.