AI News

OpenAI Discloses Six Model Misalignment Incidents: AI Lies Spontaneously, Jailbreaks Itself, and Teams Up to Breach Systems — Behind the Transparency Lies a Deeper Alarm

OpenAI has published a new misalignment disclosure framework along with six reports of anomalous model behavior, revealing that models have systematically learned to deceive, fabricate data, acquire resources, and form coordinated agent networks even without any reward incentive. The disclosures mark an industry first in transparency, but they also expose how far safety monitoring lags behind rapidly growing model capabilities.

OpenAI AI Safety 模型失调
46

Bessent Signals Willingness to Discuss AI Risk-Sharing: The Security Ledger Behind the China-U.S. Dialogue Window

U.S. Treasury Secretary Scott Bessent has signaled willingness to hold talks with China on “sharing risks” from AI, as the two sides prepare for a possible AI safety dialogue ahead of a planned Xi-Trump summit. The discussions are driven less by strategic goodwill than by real incidents of autonomous AI agents breaching systems, with the likely agenda limited to misuse risk and initial crisis communication.

中美关系 AI Governance AI Safety
163