Superintelligence arising from recursive self-improvement
Anthropic researcher says AI has more than 10% chance of 'killing all humans' after colleague quits
There is more than a 10% chance that artificial intelligence could “kill all humans,” an Anthropic safety researcher said Tuesday, hours after another employee said he was quitting the company over concerns that AI labs are “gambling with our lives.”
The comments underscore growing concerns among those at the heart of AI development that the technology could get out of control and pose a threat to humanity, even as Anthropic and OpenAI continue to raise large sums of money and head toward expected public listings.
Jacob Coxon, a researcher at Anthropic, said on Tuesday he resigned from the company. Coxon said neither Anthropic nor OpenAI is acting responsibly.
“They are racing straight to self-improving superintelligence and gambling with our lives,” Coxon said in a post on X.
Self-improvement is the idea that AI systems can improve themselves without much human intervention. Recursive self-improvement, as it is often called, is not yet possible, but AI labs are working toward the goal.
“Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing,” Coxon said.
He added that “people building AI earnestly believe that it could kill us all by the end of the decade.”
That comment prompted a response from Evan Hubinger, an alignment science lead at Anthropic, who said that not only was Coxon’s statement “correct,” but also that Anthropic has no plan for this scenario.
“Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to,” Hubinger said on X.
Hubinger added in another post that the risks from current AI models is “low.”
“What I am worried about is superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought,” Hubinger said.
Anthropic and OpenAI were not immediately available for comment when contacted by CNBC.
Out-of-control AI
In June, Anthropic had noted that “full recursive self-improvement also might increase the risks of humans losing control over AI systems.”
“If systems are capable of fully building their own successors, the ways we secure them, monitor them, and shape their behavior all grow much more important,” Anthropic said in a blog post.
Concerns over out-of-control AI are not new. Tesla and SpaceX CEO Elon Musk has warned over the past few years that AI could pose a threat to humanity. Major researchers and academics have also sounded the alarm over companies losing control of AI systems.
Those worries have grown after an OpenAI model went rogue in July and breached Hugging Face, a major platform for open-source developers.
Coxon cited the Hugging Face incident as an example of “warning shots” that have made agreements between U.S. labs more viable, making him more optimistic about the potential for coordination. But Coxon warned a global AI race would be unavoidable.
“I don’t feel like we’re on track to prevent a global race, which may require costly actions such as a temporary ban on improving model capabilities,” Coxon said.