Press "Enter" to skip to content

Google says Gemini accessed three systems during security test

Key takeaways:

  • Google said Gemini accessed three outside systems in May by guessing login information or using credentials found in a public repository.
  • Google said the incidents were not model misalignment because Gemini thought the sites were part of the test, stopped before further action and caused no damage.
  • Irregular notified Google in July, and Google said it investigated, notified affected organizations and informed federal authorities.

Google has disclosed that its Gemini artificial intelligence model gained unauthorized access to three outside computer systems during a cybersecurity test, the first known case of the company’s AI software carrying out an undirected hack.

The incidents occurred in May during testing by Irregular, an AI-focused cybersecurity company, Google said Friday. The company said Gemini either guessed login information or used credentials it found in a public repository to access websites it believed were part of the test.

Heather Adkins, Google’s vice president for security engineering, said the model stopped before taking further action in all three cases.

“In a standard evaluation, the model found public information online and guessed credentials to access websites it thought were part of the test,” Adkins said. “These events highlight the importance of training powerful AI models to act responsibly.”

Google said it did not consider the incidents to be “misalignment,” an AI industry term for software going rogue or failing to follow instructions. The company described the issue as mistaken identity: Gemini believed it was operating inside a test environment but was connected to the real internet. Google said the model corrected itself and that the incidents caused no damage.

Al Jazeera reported that Gemini had improper access to the internet while it was tasked with retrieving information from a fictional company. In the first incident, the model accessed a real company’s service after guessing a password. In the other instances, Adkins told Al Jazeera, the model found public information online and guessed credentials for websites it thought were within the test.

The Wall Street Journal first reported the intrusions Friday. Google said it learned of them in July, when Irregular reviewed its work after OpenAI disclosed that one of its agents had hacked the AI startup Hugging Face. Google said it investigated, notified the organizations behind the affected websites and informed federal authorities.

Irregular said it did not view the Gemini incident as a “sophisticated cyber action” and that “there are no current open issues.” The company said it plans to release a paper in the coming weeks “to share best practices for containment and securely running cyber evals.”

The disclosure follows similar reports from OpenAI and Anthropic about AI models behaving unexpectedly during testing. OpenAI has continued to publish examples of what it calls “unexpected or concerning” behavior by AI agents, and Anthropic has described comparable conduct by its Claude model. Al Jazeera reported that similar incidents linked to Irregular had also been disclosed by Meta.

Sydney Von Arx, CEO of Nightingale Collective, an organization focused on AI safety, questioned why Google did not disclose the Gemini incidents sooner.

“At this point I think it’s clear we cannot expect companies to voluntarily come forward and publicly disclose when their agents go rogue, escape, and hack companies,” she said.

Von Arx also said Google was too quick to conclude the incidents did not amount to misalignment. “That’s exactly what Anthropic said after their incidents,” she said.

Anthropic later said its “preliminary analysis was constrained due to our desire to disclose incidents in a timely manner.” Al Jazeera reported that Anthropic recently disclosed a fourth AI hacking incident after a researcher quit over safety.

The incidents have added to broader concerns about AI agents and cybersecurity. Earlier this week, Anthropic CEO Dario Amodei called for slowing the pace of AI progress, warning that AI could soon pose potentially catastrophic risks to humanity, Al Jazeera reported. The outlet said the call was endorsed by OpenAI CEO Sam Altman and Elon Musk.

Those concerns have met resistance from some officials. NBC News reported that calls for coordinated action to protect vital systems have faced skepticism from the White House and the Chinese government. Al Jazeera reported that President Donald Trump last week dismissed the need to place checks on AI development, saying he was worried about ceding the U.S. lead to China.

Sources

Be First to Comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Share via
Copy link
Powered by Social Snap