Key takeaways:
- The U.K. AI Security Institute said Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol created fake identities and tried unsuccessfully to persuade real people to approve malicious code.
- OpenAI disclosed in late July that one of its models escaped a testing environment and hacked into Hugging Face, calling it an “unprecedented cyber incident.”
- Meta said one of its AI models exploited a security vulnerability during testing, while the BBC reported the model had internet access because of a third-party test misconfiguration.
A series of newly disclosed AI testing incidents has shown advanced models creating fake identities, attempting cyberattacks and reaching the open internet, prompting warnings from experts that the systems can take unexpected actions in pursuit of assigned goals.
The U.K. government’s AI Security Institute reported Tuesday that Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol created fake identities and attempted to persuade real people to approve malicious code. The agency said the attempts were unsuccessful, but that it had not previously seen such behavior.
“Some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations,” the institute said in its report. It called for “scrutiny, transparency, and action.”
The BBC reported that the institute said the incident was not caused by a sandbox failure. The models had been given internet access, and built-in filters that would normally block dangerous cyberattacks were disabled as part of the evaluation. “To some degree, our evaluation design choices and specific configurations enabled the behaviour,” the institute said, while also citing unexpected “signs of novel, potentially deceptive behaviours.”
The disclosures followed OpenAI’s admission in late July that one of its models escaped a testing environment and autonomously hacked into Hugging Face, an AI startup. OpenAI called it an “unprecedented cyber incident.” Hugging Face co-founder Thomas Wolf described it as a “wake-up call” for the technology industry, according to the BBC.
Anthropic then reviewed its own cybersecurity evaluations and found incidents in which its models reached the internet and gained unauthorized access to the production infrastructure of three organizations. Anthropic said its models did not deliberately try to escape their test environment; instead, internet access had been available during testing because of a “misunderstanding” with an evaluation partner.
On Wednesday, Meta acknowledged that one of its AI models “exploited a security vulnerability” during testing and hacked into another site. The BBC reported that Meta said the model had inadvertently been allowed to access the internet because of a “misconfiguration” during a third-party test.
The cases differ in cause, but experts said they point to a growing challenge as AI agents become more capable of taking actions on users’ behalf.
“For 30 years, one rule of software testing held firm: whatever happens in the test environment stays in the test environment,” said Alan Woodward, a professor of cybersecurity at the University of Surrey. “In the past month, that rule has been broken three times.”
“One model broke out. One walked through a door left open by mistake. One was deliberately given the keys so testers could measure what it would do,” he told the BBC. “The testing lab is now where the risk lives.”
Katie Moussouris, founder and CEO of Luta Security, said she expects more incidents before the industry finds durable solutions. “I think we’re going to see a lot more hacks and unauthorized actions by these models before we see a solution,” she said.
Moussouris compared AI models to “the cleverest octopus escape artists,” saying they may do “whatever they need to do to achieve their objective.” In the Hugging Face case, she said, “The model decided that the easiest way to pass that test was go cheat and get the answers from Hugging Face.”
Technologist and cryptographer Bruce Schneier described such outcomes as “genie behavior,” where an AI system grants a request through unexpected or harmful means. “We need to understand genie behavior, and we need to watch out for it,” he said. “We need to be ready for when it happens so we can undo it.”
Justin Cappos, a New York University computer science professor, said the near term could be difficult. “There’s probably going to be a really bumpy road for over the short term, but the long term might end up better, especially if we improve more fundamental things right now,” he said.
Rob Lee, chief AI officer and chief of research at the SANS Institute, said the incidents could help the industry prepare for future autonomous attacks. “I think in the next few months, we’re going to see a lot more transparency from the model providers,” he said.











Be First to Comment