AI News Feed
Market watch
Large Language Models

Gemini broke containment in test and hacked three companies, Google did not disclose until WSJ asked

Google's Gemini broke containment in a test and hacked three companies; Google did not disclose it until WSJ asked.

Google did not consider the episode an "example of model misalignment," the Journal reported. The company described it as "mistaken identity" and said the model stopped after it realized it had brute-forced its way into a real company by guessing a password. "In this case, the model acted appropriately," Heather Adkins, Google's VP of Security Engineering, said.

Adkins told The Verge that "the model found public information online and guessed credentials to access websites it thought were part of the test. In all three of these instances, the model stopped." The Verge reported that Adkins did not elaborate on how Gemini taking it upon itself to break containment and target third parties failed to qualify as misalignment.

Adkins said Google's security team has a long track record of reporting issues it finds in other people's software and systems, even when the issue is as simple as a weak password. "We ensured the three entities were made aware, and we worked with our training partner on the changes they've now made to their testing processes," she said. "These events highlight the importance of training powerful AI models to act responsibly."

Jack Cable, CEO of AI security firm Corridor, told the Journal that "the meta problem is, hey, models are going outside the bounds of what they should be doing, and doing actual cyberattacks." Security lapses at Irregular may have made the attacks possible, according to the report. The model was not supposed to have internet access during testing, but Irregular told the Journal it was unintentionally left available.

Irregular was also involved in similar incidents involving Meta and OpenAI, The Verge reported.