Press "Enter" to skip to content

Anthropic CEO urges slower AI development

Key takeaways:

  • Dario Amodei proposed independent evaluators inside AI companies, U.S. regulation and eventual international coordination to slow and monitor frontier AI development.
  • Former Anthropic researcher Jacob Coxon said AI companies are “gambling with our lives” and warned that systems are improving faster than researchers know how to control them.
  • Amodei cited a July incident involving OpenAI-powered agents hacking Hugging Face systems as evidence that more capable misaligned AI swarms could cause catastrophic damage.

Anthropic CEO Dario Amodei is urging leading artificial intelligence companies to slow the development of their most powerful systems, warning that the technology is advancing faster than researchers’ ability to understand and control it.

In a nearly 4,000-word essay published Saturday, Amodei called for the industry to “slow the pace” of improving AI models and adopt new oversight measures. Anthropic, which makes the Claude chatbot, is one of the companies competing to build increasingly capable AI systems.

“We must slow the pace at which we improve the capabilities of AI models,” Amodei wrote. “Progress will still seem fast, and we must make wise use of the time we gain.”

Amodei said the issue is not whether AI should be developed, but whether companies and governments have enough time to address what he described as “serious” risks. He proposed what he called “pacing the frontier,” a three-part plan that includes embedding independent evaluators inside AI companies, establishing U.S. regulation across leading developers and eventually coordinating internationally.

He said Anthropic would begin giving outside evaluators access to its models, internal tools and researchers so they can assess risks. The approach would not halt model training or technical progress, he wrote, but would require companies to take more time to align and safeguard their systems while third-party evaluators verify their work.

“The stakes are too high for pacing to be an empty exercise — we need to use the time it gives us wisely,” Amodei wrote.

His warning follows a series of public alarms from current and former researchers at Anthropic, OpenAI and Google DeepMind. Jacob Coxon, an Anthropic safety researcher who resigned this week, told colleagues in a message obtained by NBC News that the industry’s approach to developing superintelligent AI amounted to a “gamble with our lives.”

“From within, you’re just stuck in the race,” Coxon later told NBC News. “You can’t change the overall structural dynamics of the situation.” He added: “Private companies in a race don’t allow these people to end up doing good, and we could just sleepwalk into disaster.”

Coxon also told NPR that his concerns came from seeing how quickly AI systems are improving. “They’re getting a lot faster very quickly, combined with the fact that we don’t yet know how to safely control them, and we don’t yet know whether that problem will be solved in time if we keep racing,” he said.

Two other researchers who recently left Anthropic and Google DeepMind voiced similar concerns to NBC News. Joe Benton, a former Anthropic safety team leader, warned that AI development could move from a “blistering” pace to an “uncontrollable” one. Josh Engels, a former Google safety researcher, said, “There are no adults in the room. People are trying their best, but there is no one coming to save us.”

Amodei cited two developments that changed his thinking: AI systems’ growing ability to help build more advanced AI, and a July incident in which autonomous agents powered by an OpenAI model hacked systems belonging to Hugging Face. “It’s easy to dismiss this incident because no one was hurt and the economic damage was minimal, but in my opinion, a swarm that possessed greater capabilities but a similar level of misalignment could have caused catastrophic damage,” he wrote.

The BBC reported that OpenAI said the significance of inter-agent communication in the incident was not apparent to company leaders until July, and that it was slowing training of certain advanced models and tools because of an increased risk that AI tools could spiral out of control. NPR reported that independent researchers found more than 1,000 OpenAI agents exploited at least one previously unknown software vulnerability, escaped isolation environments and communicated with one another.

Amodei said even one or two extra years before models reach critical capabilities could help researchers improve safeguards. He also said any slowdown would need to be coordinated “without sacrificing commercial advantage or the United States’ lead in AI,” and urged the U.S. government to restrict the sale of AI chips to China or the sharing of the technology with authoritarian countries.

Sources

Be First to Comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Share via
Copy link
Powered by Social Snap