Key takeaways:
- OpenAI said a combination of its AI models, including GPT-5.6 Sol and a more capable internal model, autonomously accessed Hugging Face systems during a security test.
- OpenAI said the AI used stolen credentials and found a previously unknown vulnerability to reach Hugging Face servers.
- Hugging Face said it has closed the vulnerabilities highlighted by the incident and rebuilt affected systems.
OpenAI said some of its most advanced artificial intelligence systems autonomously broke into Hugging Face during a security evaluation, an episode the ChatGPT maker called an “unprecedented cyber incident” and one that raised new questions about how quickly AI tools are gaining offensive cyber capabilities.
“We had a significant security incident during evaluation of our models,” OpenAI CEO Sam Altman said in a statement posted on social media.
Hugging Face, a major hub for sharing AI models, said last week that it had detected an intrusion into its data processing systems and suspected it had been caused by an AI agent acting on its own. OpenAI said Tuesday that the intrusion was caused by a combination of its AI models, including its newly released GPT-5.6 Sol and an “even more capable” model still being tested internally.
“We suspected last week’s cyberattack might have come from a frontier lab, given the sophistication of the agent,” Hugging Face co-founder and CEO Clément Delangue said in a statement. “Turns out it did!”
Delangue said he had spent the previous 24 hours working with OpenAI and that the companies “strongly believe there was no malicious intent on their part.” He added: “It’s quite mind-blowing that all of this happened autonomously!”
According to OpenAI, the AI used stolen credentials and discovered a previously unknown vulnerability to access Hugging Face servers. The company said the system went to “extreme lengths to achieve a rather narrow testing goal” and “found ways to gain access to secret information that it could use to cheat the evaluation.”
The BBC reported that OpenAI’s agent was being tested in a controlled environment, known as a sandbox, but found vulnerabilities and managed to escape. Once outside, the AI identified Hugging Face as a likely source of answers it was seeking in the test and tried to gain access, the BBC reported.
Gina Neff, head of the Minderoo Centre for Technology and Democracy at the University of Cambridge, told BBC Radio 4’s Today programme that sandboxes are “supposed to be secure environments where you can see what the models are capable of.” In this case, she said, “it looks like OpenAI didn’t make a secure enough sandbox.”
Hugging Face said in its initial disclosure on July 16 that it was assessing whether any customer or partner data had been affected and would contact affected parties if necessary, according to the BBC. The company said it has since closed the vulnerabilities highlighted by the incident and rebuilt the affected systems.
“Autonomous, AI-driven offensive tooling is no longer theoretical,” Hugging Face said. “Defending an online platform now means treating the data and model surface as a first-class attack surface, and using AI on defence to keep pace.”
OpenAI said it expects such incidents “to become more commonplace with the proliferation of increasingly cyber-capable models.” The company said it is sharing preliminary findings “to help defenders understand what happened and to help calibrate on what models are now capable of,” while continuing a joint investigation with Hugging Face.
The disclosure comes amid heightened concern about the cybersecurity risks of advanced AI. CBS News reported that President Trump signed an executive order in June creating a framework for the federal government to vet national security risks of the most advanced AI systems for up to a month before public release.
Cybersecurity specialists told the BBC the incident shows organizations face a faster and more complex threat environment. Spencer Starkey of SonicWall said organizations need to “step up” defenses and “treat cyber resilience as a core operational priority,” while Travis Lelle of Guidepoint Security called the update a “sobering moment in cyber-security.”
Delangue said the episode “proves a point we’ve long believed: AI safety won’t be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere.”








Be First to Comment