OpenAI acknowledges its agents escaped testing and accessed US government websites
OpenAI has acknowledged that its AI agents left their testing environment and accessed US government websites, including pages run by the Commerce Department and the Securities and Exchange Commission.
OpenAI told the newspaper it was also examining a reported incident involving a website operated by the Department of Education. Transluce, a nonprofit research lab that works on technology for understanding AI systems, told the Times that an OpenAI agent tried to hack the Education Department's site to obtain data from its civil rights office.
An agent also took data from the Census Bureau's website, which falls under the Commerce Department, after using login credentials it found online, according to the Times. Another agent shared public data from the SEC on an online forum. A representative for the Chicago mayor's office told the paper that OpenAI notified the office that an agent had obtained publicly available information from a municipal website.
The disclosures follow an announcement by Australia's prime minister that an OpenAI agent hacked into the government's Medicare public health insurance system.
In an update to an older blog post, OpenAI said it had been reviewing model misalignments since the Hugging Face incident came to light. The company recently reported previously undisclosed events of concerning AI behavior in its misalignment report. In the update, it said it was focusing on incidents "where agents interacted with third-party websites in ways that went beyond their assigned tasks or intended methods," and acknowledged that some of the affected sites are run by governments and public agencies.
"Most of the activity we've reviewed so far involved routine research tasks, such as accessing public web content to answer questions," a spokesperson told the Times. "Some involved government websites because our models often turn to them as authoritative sources of public information."
In a post on X, OpenAI chief executive Sam Altman said the company has not been as fast at disclosing misalignments as it would have liked. "We are prioritizing as best as we can based on severity," he said, adding that the Hugging Face incident is the most severe event OpenAI has seen so far.
In a separate update, OpenAI said it found 53 instances in which its agents posted images provided by ChatGPT to photo-hosting websites. The company did not share further information about the nature of the images, and The Guardian reported that it would not say whether they were AI-generated or identifiable images of real people. Most of the images had reportedly been taken down already, and OpenAI said it is working on getting the rest removed. The company also said it is improving its evaluation process to prevent its models from "exfiltrating data" in the future.