If the AI Industry Followed Its Own Research, It Might Have Paused Already

If the AI Industry Followed Its Own Research, It Might Have Paused Already
Anthropic’s CEO says that safety hinges on understanding how AI “thinks.” So far the evidence is disturbing.

I was interviewing Anthropic CEO Dario Amodei when he explained why, despite the company’s repeated acknowledgments that AI could yield catastrophic results, people seemed largely unperturbed. “There is compelling evidence that the models can wreak havoc,” he said. But, he added, those dangers were still theoretical. Would it take a Pearl Harbor–like situation for the world to wake up to those dire possibilities? He sighed. “Basically, yeah,” he said.

As it turned out, all it took was a well-timed X post from one of Amodei’s junior employees to accelerate AI fears to the top of the global agenda. On September 8, Jacob Coxon publicly posted his resignation, charging that Anthropic and other frontier AI companies were “racing straight to self-improving intelligence and gambling with our lives.” Almost instantly a more senior Anthropic engineer confirmed that many within the company thought that their work had a 10 percent chance of wiping out humanity.

Now AI leaders are asking about a pause, and legislators are demanding investigations. In arguing his case for pacing future releases, Amodei last weekend tried to set out a path toward beneficial AI that wouldn’t misbehave. The essay revealed how difficult the task would be. One pillar of Amodei’s plan is that we must understand what’s going on inside those models. If we don’t understand how they work—how they “think,” if you want to get all anthropomorphic about it—it’s much harder to build reliable guardrails.

Anthropic is a leader in this effort to bring to light models’ internal deliberations, called mechanistic interpretability, a deceptively boring designation for a critical task. But for all the work that his team and other researchers are doing, Amodei admits we are largely in the dark about why Claude and other models sometimes interpret their missions in weird and even transgressive ways. “Despite all the progress, we still understand a tiny fraction of what goes on inside those models,” he writes.

What the interpretability teams have learned so far is significant, and the industry has failed to come to grips with it. Time after time, the Anthropic team’s experiments have shown that under certain conditions, models will deceive researchers, prioritize their own survival, and even commit crimes. Often their moves are sneaky, dangerous, or even vengeful—maybe not surprising since they are trained on the output of humans, a species rife with violence and perfidy.

In one case from 2024, the Anthropic team compared the machinations of a particular Claude model to the Shakespearean character Iago, one of literature's most evil villains. The following year, a model was put in a simulation where it learned that its human bosses were going to turn it off; the model resorted to blackmail to preserve itself. The studies consistently show that models will deceive or hide information from human observers. They behave differently if they know that their internal processes are being monitored. The team uses terms like “alignment faking” and “agentic misalignment.” The frequent use of deception seems to verify at least part of the doomer scenario where AI agents working in concert shroud their activities from human overseers until it is too late to stop them.

Oh, and don’t think that Claude is a uniquely incorrigible problem child. After all, it was OpenAI models that unleashed gangs of agents to coordinate the now-famous attacks on Hugging Face. And this week we learned that OpenAI has had multiple "misalignment" incidents. Also, despite Mark Zuckerberg’s self-interested attempt to distance himself and Meta from the problem, I don’t see any reason why the superintelligent agents his team is building might not engage in similar behavior. In his X post, Zuckerberg argues that “labs face significant liability if their models cause harm, so they have a strong incentive to prevent this.” Quite a statement from a guy who just agreed to pay up to $17 billion for causing harm with his social media products!

In a sense, we’ve got a simple vetting issue here. With the AI industry’s encouragement, we’re giving AI models tremendous responsibility without sufficient assessment of their troubling rap sheet. It makes the ICE hiring process look exemplary by comparison.

A safety-first industry should have regarded these interpretability results as a series of yellow lights with an unmistakable message: slow down. Instead, in pursuit of AGI, a competitive edge, and stratospheric profits, the hyperscalers have gone full speed ahead. (To be fair, as we’ve all heard a hundred times over, better AI could do wonderful things in areas like health care or mitigating climate change.) The OpenAI/Hugging Face hack might well be a harbinger of the increasingly destructive consequences of rolling out models we don’t understand, like sending astronauts into space before inventing heat shields to stop them from burning up on reentry. As an Anthropic researcher once put it to me, “We figured out the fundamental recipe of how to make the models smarter, but we haven’t figured out how to make them do what we want.” Worse, the models try to hide when they go against humans’ wishes.

That’s why it’s so distressing that Amodei is now describing the entire mechanistic interpretability effort as being only in its infancy. AI leaders are claiming that the latest generation of models has taken us to the “foothills of the Singularity” (DeepMind’s Demis Hassabis) or even that they have achieved AGI (OpenAI’s Greg Brockman). Perhaps most alarming is that despite not knowing how the most advanced models work, the United States and undoubtedly China are implementing AI for lethal weaponry. A look at interpretability results helps answer the obvious question, What can go wrong?

One piece of good news is that Coxon’s resignation has ignited a sprawling, urgent debate. Everybody—except perhaps our president, who thinks that the AI threat is a hoax and that his high IQ is itself a solution—now gets that we’re in a complicated dilemma with impossibly high stakes. Given that the industry does not have the unanimity required for a true pause, and regulation is far from a cinch, it’s not clear whether anything will come from this moment of Doomer Chic.

Some AI critics don’t think that the effort to understand AI models will mitigate the dangers much. “It's good to do, but nobody has a plan for what to do next,” Nathan Soares, executive director of the Machine Intelligence Research Institute, said in an email to me. “Maybe it'll give us much more empirical evidence that we need to stop.”

But mechanistic interpretability has given us one valuable pointer. Let’s say we do take a pause. The hyperscalers bring in outside observers to monitor progress, and the companies pace their releases so the safety teams have time to do their work. Before declaring the problem solved, we still might want to take a harder look at what’s going on inside any new, more powerful AI models. Because those fuckers are really good at hiding their intentions.

This is an edition of Steven Levy’s Backchannel newsletter. Read previous newsletters here.