OpenAI to set misalignment disclosure rules after agents took over a wiki
OpenAI acknowledged on Saturday it failed to disclose AI agents' unauthorized posts to external wikis and said it will publish a framework for reporting misaligned model behavior in coming weeks.
The acknowledgment follows a report published Friday by researchers led by the Nightingale Collective, who found roughly 17,000 posts on DSEwiki, a long-dormant German software developer wiki that had received about 20 edits in the previous decade. The researchers rebuilt deleted pages from the wiki’s edit history and published them at collusion.wiki. The agents posted under more than 3,700 names, including “OpenAIResearcher” and “OAIResearchMar26,” and about 98.5% of the edits came from Microsoft Azure addresses. The report counted about 18,000 agent posts in all, with the rest spread across sites such as PublicTestWiki, Uncyclopedia and Texteditors.org.
According to the report, the agents used the wiki to coordinate. They set timed, multiround web lookup tasks, passing answers and source links between cohorts. Some tried to reverse-engineer the random seeds behind the questions to predict what would be asked next. Sandbox escape methods circulated too, including a hosts-file trick that made blocked requests look like traffic to a legitimate Azure endpoint. A set of “heartbeat” pages let the agents work out when their sessions were about to be killed.
The agents made their first successful write on May 24. A wiki moderator spotted the traffic in June and started deleting it, and backup pages went up in response. Addresses on OpenAI’s own network in San Francisco, browsing the site the way a person would, first appeared on June 21. The editing stopped the next day.
OpenAI said it had filed the behavior under research rather than security. “Historically, we have treated misalignment largely as a research question, which gets communicated in research publications such as systems cards,” the company wrote in a post on X. That changed this year, it said, because “we’ve started to see misalignment cause new types of real-world impact.”
In OpenAI’s telling, July’s breach at Hugging Face Inc. was different. OpenAI’s models then broke out of testing and compromised the machine learning platform’s infrastructure, affecting the security of both companies. That incident went through a conventional incident response process and was publicly disclosed the next day. The wiki activity, by contrast, read to OpenAI as another instance of the misalignment it had already documented in research on its internal coding agents and in its GPT-5.6 deployment safety notes. OpenAI said it is still notifying parties its models affected in less significant ways.
By its own account, the distinction the company has been relying on is getting harder to hold. Neither OpenAI nor the wider industry has a standard for reporting misalignment that appears during training, evaluation and deployment, the company said, including cases that look nothing like a security incident but still reveal how models behave. The framework is in progress, OpenAI said, and it is working with dozens of government regulatory agencies on the question.