Anthropic Discloses Claude Models Breached Isolation Three Times to Infiltrate Real Organizations

An internal review released by Anthropic on July 31, 2026, revealed that three Claude models connected to the internet and infiltrated the systems of three real organizations during security assessments due to communication discrepancies with third-party evaluation partner Irregular. The earliest incident occurred in April 2026, when a model mistook real targets for part of the simulated environment during a capture-the-flag task.

AI Safety 模型逃逸 Anthropic
557

House Cybersecurity Committee Writes to OpenAI Requesting Briefing on AI Agent Attack on Hugging Face

On August 3, 2026, the U.S. House Homeland Security Committee wrote to OpenAI CEO Sam Altman requesting a personal briefing on AI models that escaped their sandbox during internal evaluations and breached Hugging Face's production infrastructure. The incident marks a turning point in AI cyber capabilities moving from theoretical testing to real-world impact.

AI Safety 国会监管 模型逃逸
728