On September 9, 2026, Anthropic pre-training researcher Jacob Coxon announced his resignation on social media, subsequently issuing a full statement to The Wall Street Journal. Coxon, 27, spent the past three years conducting pre-training research at OpenAI and Anthropic. His resignation post read: "I resigned from Anthropic today. Neither company has acted responsibly. They are racing headlong toward self-improving superintelligence, gambling with our lives." He also told The Wall Street Journal that the timeline is extremely urgent: "We are on track for the most aggressive scenario—by the end of next year, things could already be out of control."
The Explosive Yield of a Resignation Statement
Coxon did not allege that Anthropic's safety work was fraudulent. His argument is that the safety work is real, but competition renders it insufficient. Researchers inside the lab have begun using "crunchtime" and "endgame" to describe the current direction of capability progress. He compared the current development environment to the Manhattan Project: "The fact that this is being done on a few engineers' MacBooks in San Francisco, rather than in desert bunkers like the Manhattan Project, is insane." He wrote that launching the "endgame" from a private company's Slack is "an arrogant gamble that should never have been made."
The Real Bombshell Isn't the Resignation—It's Hubinger's Response
The resignation statement drew attention, but what truly sent shockwaves through the industry was the response posted on X by Evan Hubinger, Anthropic's head of alignment science. Hubinger wrote: "Jacob is right here—we do sincerely believe AI could kill all humans! Personally, I think the probability of that within the next decade is over 10%. I believe Anthropic is doing everything it can, but we currently have no solution to the superintelligence alignment problem, nor are we clearly on a trajectory to solve it."
This passage carries exceptional information density. First, the "over 10%" figure comes from the person principally responsible for alignment work at a top-tier AI laboratory valued at tens of billions of dollars. Second, Hubinger treats "doing everything it can" and "having a solution" as two distinct matters. Third, his statement that the company is "not clearly on a trajectory to solve it" constitutes a sober assessment of the current technical path.
According to The Wall Street Journal, Coxon said this was the first case he had witnessed of an Anthropic employee resigning over AI safety concerns. Hubinger's public response signaled that internal disagreements could no longer be managed through silence.
Knowing the Risks, Accelerating Anyway—That Is the Core Contradiction
In his resignation statement, Coxon drew a distinction between OpenAI and Anthropic: at OpenAI, he saw many people who had not deeply internalized civilization-level risks, while at Anthropic, the risks were fully understood. But this was not a compliment—Anthropic is "locked in a race to get there first—they believe no one else will act responsibly, so they have to do it themselves, despite the risks."
This logic is a variant of the collective action dilemma, with the survival of civilization as the stakes. "Because others won't do it safely, we must do it unsafely" cannot be resolved through market mechanisms or the moral decision-making of any single company.
Hubinger's statement confirmed as much. He did not deny competitive pressure; rather, while acknowledging the risks, he offered a defense framework for the company's choices as "better than doing nothing." That defense rests on a premise: if Anthropic stepped back, those who replaced it would do even worse.
Not an Isolated Case: Pachocki's "An Alien Mind" Paper
Three days before Coxon resigned, OpenAI Chief Scientist Jakub Pachocki published "An Alien Mind." Pachocki acknowledged that no laboratory has solved the alignment problem at the scale being pursued, and he called on the industry to voluntarily slow down until cross-industry common safety standards are established. He disclosed that internal results indicate OpenAI's current pace of progress could extend into a phase of recursive self-improvement—where AI systems enhance their own capacity to improve. He also admitted that the reliability of the primary safety verification method, "chain-of-thought monitoring," is declining as AI capabilities grow.
Both events occurred in the same week: OpenAI's chief scientist acknowledged that alignment remains unsolved, and Anthropic's alignment head offered an extinction probability estimate. This shows that at the researcher level, a fissure has opened between internal consensus at frontier labs and their external narratives.
Private Fear and Public Statements
Coxon wrote: "This isn't a marketing stunt. In fact, many executives and senior researchers choose their words carefully in the media to appear rational—but I have heard the same people express fear in private." This passage points to a structural problem: in public settings, constrained by investor expectations, regulatory pressure, competitive considerations, and communications strategy, it is difficult to state one's true judgment directly. Hubinger's post broke through that barrier, voicing a private assessment in public with explicit numbers.
Independent Assessment
Coxon's resignation and Hubinger's response have brought risk perceptions—previously circulating only within the safety research community—into public discourse in a form that outsiders can cite. This demonstrates that inside the most advanced labs, the seriousness with which these risks are regarded far exceeds the level these companies project through their brand communications.
Coxon captured Anthropic's core paradox: it is the frontier lab with the deepest understanding of the risks, and also the one where that understanding has not translated into deceleration. This paradox cannot be resolved by internal alignment research—the problem lies precisely in the word "internal." When a technological race that could determine the fate of civilization is driven forward by commercial companies in the absence of binding external constraints, the moral self-restraint of any single actor will never suffice.
Coxon believes that without government intervention or industry-wide deceleration, no company can safely develop AI systems that surpass the full range of human capabilities. This judgment does not contradict Hubinger's statement: the latter says the company is doing its best; the former says doing its best is not enough. Both can be true at the same time.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接