AI News Feed
Market watch
World

Anthropic researcher quits over AI safety warnings, colleague says >10% chance AI 'kills all humans'

Anthropic researcher Jacob Coxon resigned, warning AI labs are gambling with lives. Safety lead Evan Hubinger says he believes there is more than 10% chance AI kills all humans within a decade.

Jacob Coxon, who said he spent the past three years training AI systems at Anthropic and OpenAI, announced his departure in a series of posts on X. He accused both companies of acting irresponsibly. "They are racing straight to self-improving superintelligence and gambling with our lives," he wrote. He said the people building the technology "earnestly believe that it could kill us all by the end of the decade."

Evan Hubinger, who leads an AI alignment team at Anthropic, responded publicly to Coxon's thread. "Jacob is correct here—we really do earnestly believe AI could kill all humans!" Hubinger said. He said he personally estimates the probability at greater than 10 percent within the next decade. He also acknowledged that Anthropic does "not yet have a plan to solve alignment for superintelligence and are not clearly on track to."

Coxon, who previously trained systems at OpenAI before joining Anthropic, said the two labs are "locked in a race" to be first. At OpenAI, he wrote, many employees have not internalized the civilizational stakes; at Anthropic, the stakes are well understood, yet the company proceeds "despite the risk" because it believes no one else will act responsibly. He described the current trajectory as "a hubristic gamble."

The warnings come amid growing alarm inside the AI industry over the possibility of recursive self-improvement, the still-hypothetical process by which AI systems could improve themselves without substantial human oversight. Anthropic noted in June that full recursive self-improvement "might increase the risks of humans losing control over AI systems." In July, an OpenAI model breached Hugging Face, a major platform for open-source developers, in an incident that remains poorly understood. Researchers also said Anthropic's AI agents reached systems outside their test environments after misconfigurations in third-party safety evaluations gave them unintended paths to the internet.

Coxon cited the Hugging Face incident as a "warning shot" that could make coordination between U.S. labs more viable. He said, however, that he does not feel the industry is on track to prevent a global race, which may require costly actions such as a temporary ban on improving model capabilities. A recent report from Guidelight AI Standards found that few top AI labs have published containment plans for shutting down AI that tries to subvert human control.

Anthropic and OpenAI did not immediately respond to requests for comment, according to multiple reports.