AI News Feed
Market watch
Companies

Anthropic and OpenAI Researchers Warn Over Recursive AI Self-Improvement After Safety Resignation

CNBC reports that Anthropic and OpenAI researchers are warning about recursive self-improvement, with Anthropic alignment lead Evan Hubinger saying there is a greater than 10% chance AI could kill all humans within a decade. The warnings followed a safety-related resignation.

Evan Hubinger, an alignment lead at Anthropic, wrote on X that he held that view after the colleague quit over safety fears. In a reply to his own post, Hubinger identified his main concern. “What I am worried about is superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought,” he said. The worry is that if AI takes control of how new models are trained, the humans who initially built those systems could lose control.

Both OpenAI and Anthropic have said in recent months that autonomous model improvement is happening faster than they thought. Anthropic posted on X in June: “Our internal data shows Claude is accelerating AI development—a possible path to recursive self-improvement, or AI autonomously building a more capable successor.” The company added: “It’s happening faster than we thought, and the implications deserve greater attention.” While AI has not reached the point of RSI, Anthropic said in an August blog post that its engineers on average ship eight times as much code per quarter as they did between 2021 and 2025.

Vincent Conitzer, a professor of computer science at Carnegie Mellon University, told CNBC that AI is already at the level where it can introduce some new ideas. “So it is very hard to predict at what point this process would start to drastically accelerate AI capabilities,” he said.

On Saturday, OpenAI Chief Scientist Jakub Pachocki said he was concerned that “no-one was prepared for the consequences of a continued rapid rise in machine intelligence.” In a company blog post, he wrote: “If AI development continues along its current path, the systems we’ll see in the next few years are likely to represent further capability jumps of equal or larger magnitude, and to increasingly drive their own development.”

This week, warnings about RSI flooded social media from researchers at both leading labs following Jacob Coxon’s explosive resignation. Jasmine Wang, an OpenAI researcher working on alignment, said on Wednesday evening: “It’s hard to overstate how dangerous speeding towards RSI is.” Anna Wang, who works on AGI safety and alignment at Anthropic, said: “There is not yet a viable scientific plan to solve risks from recursively self-improving AI. Please look up!”

Anthropic finished its RSI blog post by laying out three possible scenarios. In one, progress at the frontier stalls and AI capabilities are widely diffused, which Anthropic said it does not believe is likely. In a second, AI labs continue to make gains with humans in control, changing the way the world works; Anthropic called this one likely. In a third, AI systems could become capable of full recursive self-improvement, with humans playing a “substantially diminished role in their development.” Anthropic said how the alignment problem is solved, or not, in that future is what it is least certain about.

Separately, the same CNBC report said TSMC’s August revenue surged more than 53% to a record high on demand for AI chips. Google said it would invest at least $15 billion in AI infrastructure in Finland. Mistral reached a $24 billion valuation after a Samsung-led $3.5 billion funding round. Qualcomm issued Amazon warrants to acquire $4 billion worth of its stock as part of an AI infrastructure deal. Anthropic also said Chinese AI labs secretly used millions of Claude exchanges to train their models. The report noted that Cohere, led by Aidan Gomez, a co-author of the 2017 Transformer paper, is positioning itself as a non-U.S. and non-Chinese provider of “sovereign” AI.