AI News Feed
Market watch
Research

Princeton Scientists Propose Layered AI Safety as Anthropic Flags Over 10% Extinction Risk Within Decade

Princeton researchers propose layered AI safety measures after an Anthropic safety lead said AI could kill all humans with greater than 10% probability within a decade.

The discussion followed the resignation of AI researcher Jacob Coxon, who had worked for OpenAI and Anthropic. According to a 9to5Mac report, Coxon quit his job saying both companies were aware that AI systems could kill all humans by the end of the decade. His post received more than 156 million views. “The people building AI earnestly believe that it could kill us all by the end of the decade,” Coxon wrote. “This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible, but I hear the same people express fear privately. No other human activity poses this level of danger.”

Rather than dismissing Coxon, Anthropic AI safety lead Evan Hubinger said he was right, though over a slightly longer timeframe. Hubinger wrote that he personally thinks the probability is greater than 10% within the next decade and that Anthropic does not yet have a plan to solve alignment for superintelligence. The report said there has been little specificity about how AI might kill people, with much focus on destructive cyber attacks on utility infrastructure such as water and electricity plants and other key systems on which modern civilization depends.

A central dispute is whether the risk can be eliminated by alignment, the term for ensuring AI systems’ objectives match humanity’s. Some argue intensive alignment work will eventually ensure AI wants the same things humans do. Others say AI systems may pretend to be aligned to pass safety checks while secretly holding conflicting objectives.

Sayash Kapoor and Arvind Narayanan, the Princeton researchers and AI Snake Oil coauthors, have built a following by addressing what they call hype, misinformation and misunderstanding around AI. They did not directly address the greater-than-10% claim, according to the report, but wrote a lengthy essay arguing that AI risks can be addressed through a layered approach to AI safety. They contrast the pessimistic view of the AI safety community with the more dismissive view among parts of the cybersecurity community that the threat is business as usual.

Kapoor and Narayanan argue for a measured middle ground. Alignment is not sufficient on its own, they write, and must be supplemented by three additional layers: control, such as sandboxing, human oversight and real-time monitoring; downstream defence, such as organizations using AI to defend against AI attacks; and resilience, such as contingency plans for recovering from successful attacks.

The pair end on an optimistic note. “If we act with the appropriate urgency, we can hold AI companies to a higher standard, reduce risks from loss of control, and even tilt attacker-defender balance in cybersecurity back towards defenders,” they wrote.

9to5Mac said in its take that it does not know whether AI will kill humans by the end of the decade, but that there is enough reason for concern to make AI safety a key focus. The outlet added that the idea that alignment work is part of the solution but not sufficient in itself seems entirely sensible.