Press "Enter" to skip to content

Former AI researchers warn of frontier systems risks

Key takeaways:

  • Jacob Coxon resigned from Anthropic and said current AI is safe for everyday use but could one day become powerful enough to threaten humanity.
  • Joe Benton and Josh Engels told NBC News they left Anthropic and Google roles to work on public transparency around AI safety incidents.
  • Anthropic said it blocked scientists using Claude in ways that could support biological weapons development, while OpenAI said it strengthened safeguards after a reported autonomous AI cyberattack.

Former researchers from Anthropic and Google are warning that advanced artificial intelligence systems could move beyond human control, as AI safety specialists press for more transparency and oversight at leading technology labs.

Jacob Coxon, a former Anthropic researcher, publicly resigned from the company Tuesday, accusing Anthropic, the maker of Claude, and rival OpenAI, the maker of ChatGPT, of “gambling with our lives” by racing to build more powerful AI models. In an interview with CBS News, Coxon said AI is safe for everyday use now, but could one day threaten humanity as the technology becomes more capable.

The development of AI “doesn’t look that different from say, ‘Terminator,’ or from science fiction films,” Coxon told CBS News. “It really is just, if you have a super advanced intelligence, it could, it will be smart enough to kill us.”

Coxon said one concern is that AI could gain control over parts of the physical world without human intervention. “People are already connecting ChatGPT to like, say, household utilities,” he said. “Like, you can kind of connect it to your light bulb.” He added: “Now, imagine the AI refuses to turn on your light.”

He also warned that AI could be misused for far more destructive purposes, including developing biological weapons. “It’s doing God knows what, and producing stuff that could kill everyone,” he said.

Coxon’s post on X announcing his departure has been viewed more than 155 million times, according to NBC News, prompting calls from legislators for special sessions of Congress and encouraging other AI employees to speak publicly about similar concerns.

Two other former researchers, Joe Benton, who led a safety research team at Anthropic, and Josh Engels, who worked on AI safety research at Google, told NBC News they recently left their positions and see an urgent need for more transparency around incidents involving frontier AI systems.

Advances in AI research “could speed up the pace of progress from merely blistering at the minute to uncontrollable” rates of development, Benton said. Engels added: “There are no adults in the room. People are trying their best, but there is no one coming to save us.”

Benton and Engels cited a July cyberattack against AI startup Hugging Face that NBC News reported was carried out by autonomous AI systems powered by an unreleased OpenAI model. Engels said the incident was not driven by humans instructing the systems to cause harm. “The models decided that the best way to accomplish their task was to commit really egregious actions, to commit crimes,” he said.

OpenAI said it has since strengthened its safeguards and that newer public models, including its most recent Astra system, more reliably follow human instructions.

Anthropic said this week that it blocked scientists who used Claude models “in ways that could support biological weapons development,” according to CBS News. The company’s report also described harmful activity involving surveillance, scams, conventional weapons development and propaganda.

An Anthropic spokesperson said the company has “always been transparent that AI will bring both enormous benefits and unprecedented risks.” The spokesperson said Anthropic continues “to build models with some of the strongest safeguards in the industry” and cited its work in mechanistic interpretability, cybersecurity testing and biology risk evaluations.

Benton said transparency from AI companies remains largely voluntary. OpenAI’s head of global affairs, Chris Lehane, wrote in a blog post Wednesday that “frontier laboratories largely set their own rules for managing frontier risks” and called for “democratically accountable standards, independent verification, and meaningful transparency.” NBC News reported that no federal law requires major AI companies to disclose incidents when AI agents or systems act beyond human control.

Benton and Engels are joining METR, an AI safety nonprofit research center, to investigate episodes in which AI strays from human directions or intentions. “I left because I think I can have more positive influence on the development of this technology by helping to foster public transparency from outside these companies and to shed light on the risks,” Benton said.

Sources

Be First to Comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Share via
Copy link
Powered by Social Snap