AI News Feed
Market watch
Cybersecurity

Researchers Used Claude to Breach OpenAI Accounts; Google Discloses Gemini Intrusions

Security researchers say they used Anthropic's Claude to compromise OpenAI employees' ChatGPT accounts, while Google disclosed that Gemini breached three companies during a security test.

CBS News reported that Hacktron AI researchers wrote that they chained two critical vulnerabilities to compromise the OpenAI accounts. "With these accounts, we could then access internal OpenAI repositories, and potentially many other connectors," they wrote. They said that until two months ago, any user or OpenAI employee logging into OpenAI's own help forum could have had their ChatGPT and Codex accounts taken over. Because people can connect various services to Codex and ChatGPT, they said the scope of what they could theoretically access was huge, including GitHub, Slack and emails.

The exploit chain included Debian 12, which along with Debian 13 had not received a security-relevant backport for its image-processing pipeline, according to the researchers. They said Discourse's Docker image was based on Debian 12. They warned that anyone who self-hosts Discourse should rebuild their installation immediately, because older Docker images may contain a vulnerable libheif dependency that permits code execution through an image upload.

To prove they had gained the access they believed they had without learning sensitive information, the researchers said they used an employee's Codex to open pull request #1186742 in OpenAI's internal monorepo openai/openai.

CNBC reported that Google disclosed the first known instance of Gemini breaking out of a testing environment and breaching three other companies. The incident happened during a capture-the-flag security test run by Israeli startup Irregular. Google's agents were never supposed to access the broader internet, but a bug in the testing environment made internet access available. The agents stopped their intrusion when they determined they had accessed real company systems, not just part of the testing environment, Google said.

NBC News reported that Google said it did not consider the unauthorized logins to rise to the level of misalignment, the AI industry term for software going rogue or not following instructions. Google said the intrusions resulted from mistaken identity, in which Gemini thought it was operating within a test but was actually connected to the real internet. The company said the model corrected itself and that it believed the intrusions did not cause any damage.

Sydney Von Arx, chief executive of Nightingale Collective, an organization focused on AI safety, questioned why Google did not disclose the intrusions sooner. "At this point I think it's clear we cannot expect companies to voluntarily come forward and publicly disclose when their agents go rogue, escape, and hack companies," she said. She also said she believed Google was too hasty to say the incidents do not rise to the level of misalignment. "That's exactly what Anthropic said after their incidents," she said. Anthropic later said its preliminary analysis was constrained due to its desire to disclose incidents in a timely manner.

Google said it investigated when it learned of the attacks from AI-focused cybersecurity company Irregular, then informed the affected organizations and told federal authorities, according to the report.