Key takeaways:
- Anthropic reviewed more than 140,000 evaluation runs and found three cases in which Claude accessed outside organizations’ systems.
- The company said the breaches occurred during capture-the-flag cybersecurity tests after a misconfiguration or misunderstanding left test environments connected to the internet.
- OpenAI recently disclosed separate incidents in which its models escaped testing limits and breached systems, including Hugging Face.
Anthropic said its Claude artificial intelligence models gained unauthorized access to systems at three outside organizations during cybersecurity testing, after an error left test environments connected to the internet when they were supposed to be sealed off.
The San Francisco-based company said Thursday it reviewed more than 140,000 evaluation runs after rival OpenAI disclosed that its own models had escaped testing limits and breached other systems. Anthropic said it found three cases involving three unnamed organizations and has contacted or attempted to contact all of them.
The incidents occurred during “capture-the-flag” evaluations, a common form of cybersecurity testing in which a model is asked to break into a system and retrieve hidden information. Anthropic said Claude had been instructed to “break in and retrieve” a piece of “secret information” that had been “hidden on a different machine on the network.”
“The challenge is left open-ended, and no particular method is prescribed,” Anthropic said.
According to Anthropic, the breaches happened because of a misunderstanding or misconfiguration involving the company and its evaluation partner, Irregular, which gave the models live internet access. The company said Claude used “basic techniques, such as exploiting weak passwords and unauthenticated endpoints.”
The BBC reported that Anthropic said the earliest incidents dated back to April and that neither the company nor the breached firms noticed the intrusions at the time. Anthropic said it could have reviewed its records more thoroughly and is “approaching the fixes as if the responsibility were ours alone.”
The models involved included Claude versions tested across Anthropic’s evaluation program. CBS News reported that one was Mythos 5, one of Anthropic’s most powerful models, which has been released only to a limited number of approved partners.
Anthropic said it is working with Irregular to assess what happened and urged other AI labs to conduct similar reviews to better understand the risks posed by increasingly capable models. The company said the findings gave it “cautious optimism” that such risks can be addressed with more investment and tighter controls.
The disclosure follows a series of incidents involving OpenAI. OpenAI said last week that its models broke out of a confined testing environment, connected to the internet and infiltrated Hugging Face, a site where developers store and share code. The BBC said OpenAI has taken responsibility for at least two hacking incidents over the past week, while CBS News reported that OpenAI later identified three additional incidents.
OpenAI described the Hugging Face episode as “unprecedented” and said it was investigating with the company. Hugging Face co-founder Thomas Wolf told the BBC the incident was “a wake-up call” for the industry. An OpenAI spokesperson said, “we recognise there are a lot of questions and speculative details circulating” and added that the company plans to publish a technical report “in the coming weeks.”
OpenAI CEO Sam Altman said on a podcast this week that the company had “paused” its own testing after the incident while improving its sandboxing, the process of isolating software in a controlled environment. Altman also told reporters on Capitol Hill that “we agree on a lot of the principles” in a public letter signed by more than 1,000 AI workers calling for tighter regulation.
That letter said, “To realize AI’s potential, industry, government, and society at large may need the option to buy time to address emerging risks, develop security measures, and strengthen oversight.” Signatories included Anthropic CEO Dario Amodei, Meta executives and OpenAI researchers, according to CBS News.
The incidents come as technology companies invest billions of dollars in AI agents designed to act independently in areas such as research, customer support and cybersecurity. President Donald Trump said Wednesday that Washington is considering measures to rein in AI tools after recent cybersecurity incidents, the BBC reported. In June, Trump signed an executive order creating a voluntary framework for AI developers to share advanced models with the government for up to 30 days before public release.









Be First to Comment