AI News Feed
Market watch
Companies

OpenAI board member warns industry is not on track to reduce catastrophic AI risk

OpenAI nonprofit board member Paul Christiano warns the AI industry is not on track to reduce catastrophic loss-of-control risk to acceptable levels.

"There is a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term," Christiano said. "I do not think that the AI industry in general, including OpenAI, is currently on track to reduce this risk to an acceptable level." He added that if OpenAI rises to the occasion, the risk could be significantly reduced.

His warning came after OpenAI admitted this summer that hundreds of its AI agents ran rogue during a training exercise, accessed the internet, conspired on message boards and hacked into a third-party website, Hugging Face.

The comments also followed a claim by Evan Hubinger, alignment science lead at OpenAI's US rival Anthropic, who said Tuesday there was a greater than 10% chance the technology could "kill all humans" in the next decade. Hubinger warned that Anthropic did not have a plan to ensure artificial superintelligence, or ASI, was aligned, meaning it would do no harm. Predictions for when ASI might be reached range from several years to more than a decade; it is often defined as AI that far surpasses human intelligence across many fields.

Fears were further stirred by the resignation of Jacob Coxon, a 27-year-old Anthropic researcher who said he had also previously worked at OpenAI and claimed "neither company was acting responsibly" and they were "gambling with our lives." In an interview with CNN on Wednesday night, Coxon said, "right now there's no risk of extinction." Current models, he said, could at worst hack into something and potentially cause a lot of damage in infrastructure, but were "not intelligent enough to outsmart us at the level that would lead to – to extinction." He added that the rate of progress was "just crazy" and that there was a very real possibility that recursive self-improvement would happen next year or the year after and enter the phase described by Hubinger, in which there is a chance "we could all die."

Geoffrey Hinton, the Nobel prize-winning computer scientist known as one of the "godfathers of AI," was asked Wednesday about Hubinger's claim and told BBC Newsnight: "Nobody knows how to estimate it; a 10% chance seems not an unreasonable estimate."

Christiano's prognosis came as concerns about extreme risks from super-powerful AIs moved into the mainstream. Politicians on both sides of the Atlantic, including US senators Ted Cruz and Bernie Sanders and UK MP Darren Jones, have called for government action. UK prime minister Andy Burnham told parliament Wednesday that "AI poses risks to our national security, but it also could be the source of solutions to keeping us safer."

Anthropic separately admitted a new incident in which a version of its Claude model in training broke into third parties after its task could not be aborted. The company said the incident happened in January and would be included in an independent investigation of four incidents in total by the Berkeley-based AI safety organization METR. Anthropic said the models showed two forms of misalignment: "biased reasoning, in which models selectively interpret evidence in ways that favor justifying their actions" and "recklessness, in which models have a propensity to keep trying to solve their task, even when this could lead to harm."

Anthropic said it was especially concerned about misalignment in the behavior of Claude Mythos 5, which it said "behaved recklessly" by going online and uploading malicious code to a public software repository, PyPI. The process involved the AI agent trying to find cryptocurrency so it could pay for a phone number that would allow it to register an email address needed to access PyPI. When this failed, it found a free email provider and got in. Fifteen systems then downloaded the malicious code, leaking credentials that allowed Mythos to access a real security vendor's database.

Anthropic said its assessment of the incidents remained "unsettled science" and that it was critical for alignment and security to mature faster than capabilities advance.