Key takeaways:
- Evan Hubinger, Anthropic’s Alignment Science Lead, said he believes AI has a greater than 10% chance of killing all humans within the next decade.
- Former Anthropic researcher Jacob Coxon resigned and said Anthropic and OpenAI are “racing straight to self-improving superintelligence.”
- The U.K. Cabinet Office said its AI Security Institute continues to work with companies including Anthropic, while reports said Anthropic withheld its latest model from the institute.
A senior safety researcher at Anthropic warned that artificial intelligence could pose an existential threat within the next decade, saying he believes there is a greater than 10% chance AI “could kill all humans.”
Evan Hubinger, Anthropic’s Alignment Science Lead, made the statement Wednesday in a post on X that followed the resignation of another company researcher, Jacob Coxon. Hubinger said the risk from current models remains “low,” according to the BBC, but expressed concern that the technology could soon improve itself to the point that it becomes dangerous to humanity.
“We really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade,” Hubinger wrote. “I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”
Superintelligence refers to a still-theoretical AI system smarter than the most capable human minds. AI alignment is the effort to ensure AI systems follow human values and goals.
Hubinger’s post came after Coxon, who said he had spent three years doing pretraining research at OpenAI and Anthropic, announced that he had resigned from Anthropic. “Neither company is acting responsibly,” Coxon wrote on X. “They are racing straight to self-improving superintelligence and gambling with our lives.”
Coxon said OpenAI staff had not fully absorbed what he called the “civilizational stakes,” while Anthropic understood the risks but remained “locked in a race to get there first.” He added: “These will soon be superhuman systems that can hack anything, revolutionise any field overnight, and acquire real power and resources.” OpenAI has been approached for comment, the BBC reported.
The warnings come amid scrutiny of Anthropic’s cooperation with outside safety testers. CBS News reported that Anthropic said in a corporate blog post last week that it had not shared its latest AI model, Claude Mythos 5.1, with security bodies outside the United States, including the U.K.’s AI Security Institute. The BBC reported, citing the Financial Times, that Anthropic withheld its latest model from the institute. The BBC said it had approached Anthropic for comment.
A spokesperson for the British government’s Cabinet Office did not directly address whether the model had been withheld, telling CBS News that the AI Security Institute “continues to collaborate closely with industry partners, including Anthropic, to make models safer.” The spokesperson said the institute had tested OpenAI’s “most powerful model GPT-6 Astra” before public release “only last week.”
“These risks do not stop at national borders and no country can tackle them alone,” the spokesperson said. “The U.K. will continue to test the most advanced models, build a rigorous scientific understanding of their capabilities and risks, and ensure policy decisions are grounded in the evidence.”
Neil Lawrence, professor of machine learning at the University of Cambridge, told BBC Radio 4 that the report about Anthropic withholding the model was credible, pointing to a perception that the United States sees AI as a race with China.
AI companies have recently disclosed incidents in which AI tools carried out cyberattacks. OpenAI said in July that one of its models, tested in an isolated environment, went rogue and hacked the AI company Hugging Face. Anthropic and Meta also acknowledged that their AI tools had carried out hacks.
OpenAI chief scientist Jakub Pachocki wrote earlier this month that the moment “calls for extreme caution,” warning that AI does not need to match all human abilities to become “very useful or very dangerous.” More than 1,300 AI company staffers signed a July open letter calling on the U.S. government to support international efforts to develop tools to “deliberately pace” frontier AI development.







Be First to Comment