Anthropic Safety Researcher Says More Than 10% Chance AI 'Could Kill All Humans'
Anthropic safety researcher Evan Hubinger said on X that there is a greater than 10% chance AI could kill all humans within the next decade, calling the risk from current models low but warning the technology may soon improve itself.
In a post on X, Hubinger said the risk from the models that currently exist was "low," but that he was "worried" the technology might soon develop and improve itself to the point where it posed an existential risk to humanity. He did not spell out how he thought AI systems could in the future result in humans being wiped out.
The BBC described his comments as the latest in a series of increasingly stark warnings about AI, saying the debate has shifted from whether the technology truly poses a risk to how big that risk is.
Hubinger's post was written in response to another post on X from Jacob Coxon, an AI researcher who has just quit Anthropic and previously worked at OpenAI. Coxon wrote that "neither company is acting responsibly," adding that "these will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources."
The report did not say whether Anthropic or OpenAI responded to Coxon's criticism. Hubinger is described as a top safety researcher at Anthropic, a company that works on frontier AI systems and whose staff have publicly debated the pace and risk of the technology; Coxon's departure adds him to a group of researchers who have left major AI labs while voicing concern about how the work is being conducted.
Neither Hubinger's percentage estimate nor Coxon's assessment was accompanied in the report by a technical account of how such outcomes would occur, and the BBC did not attribute the risk figures to any study or internal review. The exchange appeared on X, where both researchers posted, and was reported by the BBC and circulated to readers of Slashdot, which described the warning as part of a broader shift in the AI safety debate.