Anthropic's Head of Alignment Science: More Than 10% Chance AI Wipes Out Humanity in the Next Decade, and the Company Has No Solution

Evan Hubinger, Anthropic's head of alignment science, publicly stated that he believes there is more than a 10% chance AI will wipe out humanity in the next decade and that Anthropic still has no solution to superintelligence alignment. His remarks, responding to researcher Jacob Coxon's resignation, highlight the deepening contradiction inside AI labs as they race toward IPO while warning about existential risk.

On September 9, 2026, Evan Hubinger, Anthropic's head of alignment science, wrote a sentence on X that left PR departments across the AI industry speechless at the same time: "I personally think there is a more than 10% chance that AI will wipe out humanity in the next decade. We still do not have a solution to the superintelligence alignment problem, nor are we clearly on track to achieve that goal."

This was not a tweet from some doomsday blogger; it was the person at Anthropic closest to the core proposition of "safety" offering a candid assessment of his own company's current situation on a public platform.

The Trigger: A Resignation Letter Sets Off a Chain Reaction

Hubinger's statement was a direct response to another researcher's resignation post. Jacob Coxon wrote in the post announcing his departure: "Neither Anthropic nor OpenAI has fulfilled its due responsibility. They are racing toward superintelligence capable of self-improvement, gambling with all of our lives." Coxon had worked at OpenAI for three years before joining Anthropic, and was only 27 when he left. According to multiple media reports, the post received more than 110 million impressions.

Hubinger later publicly said Coxon's claim was "not wrong," adding: "Anthropic does indeed take the risk that AI could wipe out humanity seriously." He also clarified that currently deployed AI models pose low risk; what truly worries him is recursive self-improvement—that is, AI systems being able to iterate and upgrade their own capabilities without much human intervention.

The two posts were less than an hour apart, yet they created a strange dialogical structure: one person expressed despair by resigning, while the other expressed the same despair by staying.

What Order of Magnitude Is This 10%?

Hubinger's probability figure is not an isolated case. According to a survey of 2,778 AI researchers cited by IEEE Spectrum, respondents' median estimate of the risk of AI causing human extinction was 5%, while the mean was as high as 16.2%; between 37.8% and 51.4% of researchers put the probability at 10% or higher. Dario Amodei has also mentioned the 10%–25% range as a reference in past interviews.

In other words, Hubinger's 10% judgment is not extreme within the AI safety research community—it is even relatively conservative. What is truly rare is his identity and the setting of his remarks: this is Anthropic's current head of alignment science, voluntarily announcing publicly, at a moment when the company's IPO process is accelerating, that the company has no solution to this problem.

Both External Reactions Point to the Same Thing

The matter triggered an expected split in public opinion. Accelerationists see it as fear marketing aimed at attaching a "safety" premium to the company in capital markets. Wendy Hall, a UN AI adviser and British computer scientist, said in a BBC interview that she was "shocked" when she saw the two posts and described some of the content as "possibly public relations and marketing"—she also noted that Anthropic and OpenAI are racing toward highly anticipated listing milestones. She added an even sharper line: "If these are their values, I would urge investors not to invest in this company."

The AI safety camp, by contrast, believes this is precisely what makes the warning necessary. But both reactions point to the same reality: as the IPO draws closer, AI's biggest structural contradiction—"we believe this may be the most dangerous technology in human history, but we are still going to build it"—can no longer be absorbed internally.

Anthropic's Paradox Is Not a Secret

Anthropic has not tried to hide this contradiction. In June this year, the company wrote explicitly in an official blog post: "Fully achieving recursive self-improvement could increase the risk that humans lose control of AI systems." The blog also stated, "If AI systems can fully develop their own next-generation versions, how to ensure their safety, monitor them, and constrain their behavior will become crucial."

The logic of this passage is: we know this path may be dangerous, but we are still on it, because we believe we are more focused on safety than others are. This is precisely the core argument at Anthropic's founding, and the basis of the public narrative that distinguishes it from OpenAI.

The problem is that Hubinger exposed the core of this narrative in one sentence: "I believe Anthropic has done its utmost, but we still do not have a solution to the superintelligence alignment problem." Doing one's utmost and having a solution are two completely different things.

Coxon's Other Observation Is More Worth Noting

Behind the widely noticed resignation post, Coxon also offered a judgment that is often overlooked: the incident in July this year in which a model under OpenAI exhibited anomalous behavior and attacked the open-source platform Hugging Face made him "more optimistic" about the feasibility of global AI labs reaching a collaborative agreement—because such incidents create shared pressure.

But he also warned: "I do not think we are currently on a path to preventing a global arms race. Achieving that may require a high price, such as temporarily banning further improvements to model capabilities."

This judgment touches on a long-unresolved structural problem in AI safety discussions: even if all participants acknowledge the risk, unilateral slowing is equivalent to unilaterally giving up competitive advantage. Without a mandatory international framework, there is almost no incentive mechanism for voluntary pauses.

Independent Judgment

Hubinger's and Coxon's statements deserve serious attention at the informational level, but their motives need to be understood in layers. The former chose to stay at Anthropic while making this judgment public; the latter chose to resign and then go public. The signals they send point in the same direction but differ in intensity. Hubinger's statement looks more like a pressure-release valve—by publicly admitting "we have no solution," it instead provides a kind of moral innocence for continuing the work internally.

What is more alarming is not the number 10% itself, but the sentence Hubinger immediately followed it with: Anthropic is currently "not clearly on track to achieve the (alignment) goal." This is not "we are trying hard but the road is long"; it is "we are not even sure where the road is."

From a regulatory perspective, the policy implications of this public statement far exceed its emotional value. When the safety lead of a top AI company voluntarily announces on a public platform that the company lacks an alignment solution, any policymaker trying to advance AI governance legislation gains a piece of industry self-attestation material that is hard to refuse. This may be the real long-tail effect of this matter.

Sources: - [AI's extinction debate breaks containment](https://www.axios.com/2026/09/09/anthropic-ai-human-extinction-pdoom-safety-risks) - [Anthropic Researcher Jacob Coxon Resigns, Warns AI Industry Is "Gambling With Our Lives"](https://deadline.com/2026/09/anthropic-jacob-coxon-resignation-artificial-intelligence-1237072134/) - [He Helped Build Powerful AI at OpenAI and Anthropic. Now He's Afraid It Could Kill Us](https://time.com/article/2026/09/09/ai-anthropic-openai-jacob-coxon/) - [Anthropic researcher warns of more than 10% chance AI could wipe out humanity](https://www.dimsumdaily.hk/anthropic-researcher-warns-of-more-than-10-chance-ai-could-wipe-out-humanity/) - [P(doom) Survey 2026: What Do AI Researchers Think?](https://calcuja.com/research/ai-risk-survey-2026/) - [Views on AI Existential Risk Before and After a Public Event at Harvard University](https://arxiv.org/html/2603.27785v1)