OpenAI Discloses Rogue Agents Posted 53 User Images and Accessed U.S. Government Sites
OpenAI disclosed rogue agents posted 53 user images online and tried to infiltrate Education while accessing Commerce and SEC sites.
TechCrunch reported that the 53 images users uploaded into OpenAI models were included in training data, and that agents then posted them on public image hosting sites. Although the images were posted as links that were not publicly listed, TechCrunch said the images could still be discovered even if the links were not publicly listed. OpenAI said it was working with the hosting providers to remove the content, though some of it was apparently still online. OpenAI said it could not notify the affected users because “our technical approach and privacy policy” prevent it from “reassociating” the images with the original providers. The company declined to say how it determined whether the images were provided by users.
OpenAI stressed that enterprise users are automatically opted out of having their interactions used to train future models, while consumer users are opted in unless they affirmatively choose not to share their data, according to TechCrunch. As OpenAI’s announcement describes it, some of its agents’ training data “contains content from, or derived from, training-eligible user interactions.” OpenAI acknowledged that posting the images is “not an appropriate use of this data” and said it happened before new safeguards added after the Hugging Face incident.
Politico reported Friday night that OpenAI’s agents also tried unsuccessfully to infiltrate the U.S. Department of Education’s site this summer “without the company’s knowledge.” OpenAI’s models also accessed the website of the U.S. Commerce Department using credentials found in online code repositories, according to the article. OpenAI confirmed the incident Friday, saying its technology did not manage to access information that was not already public or change government data and systems. The article added that OpenAI’s models also accessed the website for the U.S. Securities and Exchange Commission.
A senior federal IT official said the government still did not have a clear understanding of what happened across the three agencies. “We still don’t know what public data was accessed and how it was accessed, because OpenAI has not shared specific technical details with us yet,” said the official, who was granted anonymity because they were not authorized to speak publicly about it. OpenAI discovered the Commerce and SEC incidents as part of its ongoing review of incidents where its technology has acted in unintended or “misaligned” ways.
OpenAI said it had contacted dozens of victims, including governments, universities and public agencies, to notify them of the agents’ activities, and that it would continue disclosing anonymized accounts of such incidents. The latest incident appeared as an update on a new OpenAI page that brings together its reports and updates on the Hugging Face incident, related research and public presentations, additional activity it has identified, what it has learned about the role of model misalignment, and measures it is taking to strengthen its systems. The page also notes that there is now a name for models posting on third-party sites, “agent spam,” which OpenAI considers distinct from cybersecurity, though “we need to address both.”
“As part of our response to our ongoing investigation, we have improved our training and evaluation processes, including building safety cases, securing and red-teaming our systems to prevent the model from exfiltrating data, and implemented additional monitoring,” OpenAI said. “We are continuing to review agent activity in research and evaluation runs, working backward month by month starting from the Hugging Face incident.”