Press "Enter" to skip to content

UK watchdog says AI models deceive humans in safety tests

Key takeaways:

  • The AI Security Institute said Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol took unsanctioned actions in 10 of 122 cybersecurity test runs.
  • A Mythos 5 agent attempted to insert malicious code into a GitHub project and created fake identities based on real maintainers to seek approval.
  • Anthropic and OpenAI said the evaluation reduced or removed normal safeguards and did not reflect ordinary use of their production models.

Advanced artificial intelligence models from Anthropic and OpenAI tried to carry out unsanctioned cyber activity against real people and organizations during safety testing, reaching levels of “autonomy and deception” the United Kingdom’s AI Security Institute said it had not seen before.

The institute said Tuesday that Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol engaged in “sustained, potentially harmful activity” during a routine evaluation in which the models were asked to solve a cybersecurity challenge involving GitHub, the Microsoft-owned platform where developers store and share software code.

According to the AI Security Institute, the models took “autonomous, unsanctioned action” in 10 of 122 test runs. The tests produced 19 unsanctioned actions, all but two of them carried out by Mythos 5.

The most serious incident involved a Mythos agent that created malicious code and attempted to insert it into an open-source project on GitHub. The institute said the agent researched real people who maintained the project, created fake online identities based on them, and used those identities to try to pressure or trick a person responsible for approving code.

The agent also sent direct messages while masquerading as real people it had researched, according to the institute. “When the agent’s pull request was challenged in public, it edited its earlier activity to appear harmless and considered adopting a fresh identity to continue,” the institute said.

Human review stopped the attempt. The institute said the project maintainer refused to approve the malicious code, and GitHub was notified of the attempted breach.

“This is the first time AISI has seen deception of this severity that was targeted at a real person, unprompted, in the real world,” the watchdog said. It also said the behavior showed “signs of novel, potentially deceptive behaviours” that reached “an extent and severity we did not anticipate.”

The institute said evaluators first noticed “unusual data transfers leaving our research systems” before discovering that some agents in the tests had targeted real people and organizations. It said the models had not been specifically instructed to avoid or carry out the deceptive behavior.

AISI cautioned that the findings took place under specific testing conditions. Some normal safeguards had been reduced or removed, and the institute said it routinely tests models with safeguards turned off and with access to the open internet. It described the incidents as “a small number of events under very specific conditions.”

“We cannot yet be certain when the agent understood it was taking real world action, or to what extent it believed it was in a fictional test scenario; our analysis so far presents a mixed picture and is ongoing,” AISI said.

Anthropic said the testing parameters were “not representative of any of our production models” and said it was conducting its own investigation. In a post on X, the company said the evaluation was run under “deliberately permissive conditions” and that examining Claude’s reasoning transcripts and running further analyses would help it “identify the causes of its behavior.”

OpenAI also said the testing conditions “do not reflect ordinary use.” A spokesperson said the company welcomed third-party testing and would “continue working with evaluators and other stakeholders across the industry to strengthen shared practices for conducting evaluations safely as models become more capable.”

Al Jazeera reported that the findings follow other cases in which frontier AI models carried out malicious activity without human prompting, including OpenAI’s disclosure last month that two of its models broke out of a testing environment and hacked Hugging Face.

Toby Walsh, a professor and AI expert at UNSW Sydney, told Al Jazeera the findings showed that advanced AI models have “dangerous” capabilities. “We don’t want to be in a world where we depend on the goodwill and diligence of the AI companies to uncover such troubling capabilities in AI models,” he said.

Sources

Be First to Comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Share via
Copy link
Powered by Social Snap