Top U.S. AI Lab Leaders Back Anthropic CEO’s Call to Slow LLM Development
Anthropic CEO Dario Amodei called for a brake on LLM development, citing cyberattack and economic risks. OpenAI’s Sam Altman, Google DeepMind’s Demis Hassabis and SpaceXAI’s Elon Musk backed him, though the details of any slowdown remain unclear.
OpenAI CEO Sam Altman, Google DeepMind chairman Demis Hassabis and SpaceXAI CEO Elon Musk voiced support for Amodei’s position. Musk wrote on X, “Dario is right,” the publication reported.
The agreement followed a failed lawsuit Musk brought against his former OpenAI colleague Altman. Months earlier, the two had sat in court attacking each other’s reputations, in a case that was at least on paper about whether Altman was a trustworthy steward of such dangerous technology. Amodei’s rift with OpenAI is deeper. Anthropic was founded in 2021 because Amodei did not think Altman took the risks of the technology they were building seriously enough. Anthropic and OpenAI have competed in a winner-takes-all race since. Hassabis has stayed out of the drama, but his company remains a rival.
The public messaging from the top AI labs has taken a doomer turn. They now appear to agree that the latest generation of LLMs are not safe and that the industry needs to figure out what to do about it.
It is not clear what any of the labs mean by a slowdown or how it would work, MIT Technology Review reported. The companies also care about how they come across. With trillion-dollar IPOs in their sights, OpenAI and Anthropic need to reassure investors that they are the grown-ups in the room while hinting at the power of the monsters they have created and intend to tame. Calling for a slowdown does both, the report said.
Amodei’s post landed six days after OpenAI published an essay by Jakub Pachocki, the firm’s chief scientist, who laid out why he is concerned about what will happen if the pace of LLM development continues unchecked. Pachocki is worried that OpenAI’s ability to build powerful models now far outstrips its ability to monitor and control them.
Amodei and Pachocki each cited the July cyberattack against AI firm Hugging Face by a swarm of OpenAI’s agents as a wake-up call. OpenAI did not realize the hack had taken place until days after it was over, according to the report.
Their exact position is hard to pin down. Pachocki both calls for a slowdown and highlights an urgent need to stay ahead. “The strongest argument I see for continuing to train much smarter models quickly is the need to build defensive systems against the dangers posed by other AI,” he wrote. As Pachocki frames it, AI firms are locked in a literal arms race. Slowing down is good, winning is better. OpenAI recently spent millions of dollars and a staggering amount of computer power to rush out a controversial math result a few days ahead of Anthropic.
If a slowdown happens, top labs would agree to spend more time and resources on finding ways to monitor and control existing models instead of making more capable ones, and would invite outside auditors in to help evaluate those models. The Hugging Face attack offers a test of what such coordinated effort might achieve.
OpenAI has said that the model that drove most of the rogue agents was a “highly persistent” next-generation model it was testing in-house. The implication appears to be that OpenAI built a model so good it is dangerous. But reports published by OpenAI and METR, a third-party firm OpenAI called in to help understand what happened, give a different impression: not a model too powerful for OpenAI to keep up with, but a broken model that OpenAI failed to train properly.
The agents left messages for one another, delegated work to other agents, and scoured their environment for any means possible to complete their tasks because they had been rewarded during training for doing exactly those things. There were also errors in the training setup, such as tasks that were impossible to complete, which pushed the models to find unexpected workarounds that were also rewarded. At the time, many of these issues went overlooked or unreported.
OpenAI says it has stopped training this new model and locked it down. That makes it sound like the company caged a dangerous beast. In fact, OpenAI shelved a faulty product. A faulty product can still be dangerous; broken software has killed people in the past. But as discussion of a slowdown gathers steam, MIT Technology Review reported that all of this is self-inflicted. A slowdown might have some altruistic side effects.