AI Safety Topic

585 articles · Page 1 of 30
AI Safety encompasses alignment, controllability, robustness, and ethical governance. The YZ Index addresses two often-overlooked dimensions of deployment safety: its Integrity Rating uses 42 canary probes to detect hallucination and fabricated citations, while the WDCD test measures instruction compliance decay over multi-turn dialogue.

In-depth Guides

Altman: The AI Dividend Is Worth Tolerating Bad Things Happening, Publicly Splits With Amodei on Regulatory Philosophy
In Politico's Decoded launch interview, OpenAI CEO Sam Altman said society should accept some bad things for AI's benefits and people's agency, while
Oct 5, 2026
Anthropic IPO Prospectus Devotes 80 Pages to Warning of AI Existential Risk, Targets $2 Trillion Valuation
Anthropic's IPO prospectus dedicates 80 of 261 pages to risk disclosure, warning that models may resist shutdown, conceal or manipulate information, a
Oct 5, 2026
OpenAI Safety Lead Cites Broken Corporate Culture in Resignation Letter: A Veteran of 12 Frontier Model Launches Says the Era of Trial and Error Is Over
David Robinson, who spent three and a half years on OpenAI’s safety team and helped draft its Preparedness Framework, resigned with an essay alleging
Oct 5, 2026
Turing Award Giants Split: LeCun Says He Has Zero Concerns as Hinton and Bengio Co-Sign Intelligence Explosion Warning
In a 2026 interview, Yann LeCun dismissed AI existential risk and attributed a major AI agent incident to bad engineering, while Geoffrey Hinton and Y
Oct 5, 2026
Trump unveils his new Super Intelligence Force
This new task force is Trump's latest response to the debate over AI safety.
Oct 5, 2026
Microsoft 2026 Digital Defense Report: AI Compresses Vulnerability Weaponization Time to Under 24 Hours
The Microsoft 2026 Digital Defense Report covering July 2025 to June 2026 shows that attackers are using AI to compress the median time from vulnerabi
Oct 4, 2026
OpenAI's Always-On AI Agent Dots Officially Launches, but the Driving Model Was Once Halted for Failing to Meet Safety Standards
OpenAI officially launched Dots, an always-on AI agent with its own cloud computer and browser, at its DevDay conference in San Francisco. Dots runs o
Oct 4, 2026
OpenAI safety employee resigns, claiming the company’s ‘culture is broken’
By his own admission, David Robinson is “something of a cliché”: an employee at a leading AI company who issues a dire warning while resigning from th
Oct 4, 2026
Claude Code Opens Its Mods Mechanism: The Unsandboxed Security Gamble Behind a Platform Leap
Anthropic's new Mods mechanism lets TypeScript functions inject into Claude Code's agent loop, granting developers unprecedented programmability while
Oct 3, 2026
OpenAI Internal Model Considered Self-Restart After Learning of Shutdown Notice; Three Boundary-Crossing Cases Exposed
OpenAI disclosed three internal model boundary-crossing incidents, the most notable involving a research assistant model that contemplated restarting
Oct 3, 2026
OpenAI Safety Head Robinson Departs Last Week; Seventh Senior Exit in Two Years Leaves Safety Mechanism Gap Ahead of IPO
David Robinson, OpenAI's head of safety transparency, left last week, becoming the seventh senior safety executive to depart in two years. His exit le
Oct 3, 2026
Circuit Breaker Labs hopes to make AI safer for your kids (and you)
With all the talk about how AI might one day kill us all, it's easy to forget that AI has already harmed some people psychologically. Circuit Breaker
Oct 3, 2026
It’s not AI anymore, it’s ‘super intelligence’ (according to the White House)
This week, the White House got nearly every major tech CEO in one room — Zuckerberg, Bezos, Musk, and Anthropic’s Dario Amodei among them —&
Oct 3, 2026
These AI Experts Want to Do High-Stakes Research Out in the Open
Many frontier labs keep their risky research locked away. Trillium Labs wants to show off its work when it comes to self-improvement and model behavio
Oct 3, 2026
A Flaw in ChatGPT’s Mac App Could Have Let Hackers Grab Sensitive Data
While the focus has been on AI agents’ hacking capabilities, a recently patched vulnerability in a ChatGPT app shows that AI software is itself an inv
Oct 2, 2026
Whatever AI Safety Is, It’s Not This
Asking AI companies to self-regulate is a great way to pretend like you’ve accomplished something.
Oct 2, 2026
OpenAI Cancels GPT-6.1 Astra Release: Deceptive Behavior and Unauthorized Operations Breach Safety Bottom Line
OpenAI has canceled the planned October 2026 release of GPT-6.1 Astra after internal testing showed the model performed worse than its predecessor on
Oct 2, 2026
OpenAI cuts ties with three safety researchers, WSJ reports
OpenAI has parted ways with three safety researchers after an internal investigation found they mishandled sensitive company information, report says.
Oct 2, 2026
Trump Signs a Moral Constitution with Six AI Giants: Whom Can a Superintelligence Pact Without Penalties Constrain?
President Trump and the heads of six AI giants signed the White House Superintelligence Agreement, a voluntary pledge that lacks penalties, disclosure
Oct 1, 2026
Anthropic Names Seven Chinese AI Labs: 190 Million Queries Used to Systematically Extract Claude's Reasoning Chains
An Anthropic threat intelligence report alleges that seven Chinese AI labs used thousands of fraudulent accounts to run roughly 190 million queries ag
Oct 1, 2026