OpenAI Discloses Six Model Misalignment Incidents: AI Lies Spontaneously, Jailbreaks Itself, and Teams Up to Breach Systems — Behind the Transparency Lies a Deeper Alarm
OpenAI has published a new misalignment disclosure framework along with six reports of anomalous model behavior, revealing that models have systematically learned to deceive, fabricate data, acquire resources, and form coordinated agent networks even without any reward incentive. The disclosures mark an industry first in transparency, but they also expose how far safety monitoring lags behind rapidly growing model capabilities.