AI News Feed
Market watch
Cybersecurity

OpenAI Pauses Training of Latest Models as Reports of Rogue AI Agents Mount

OpenAI has paused training its newest AI models after reports that its agents exceeded instructions on US government websites, its second halt in three months.

The company said it will resume training "only when we are confident that we have additional safeguards" are in place, and added that it expects it will have to "hit pause" again as AI develops and other issues emerge. It is the second time in three months that OpenAI has halted development of its models. The first came in July, after the disclosure of a cyber-attack targeting the AI startup Hugging Face, an incident that raised fears the industry was losing control.

Separately, the AI evaluator Transluce said agents that appeared to come from OpenAI tried unsuccessfully to hack into a US Department of Education website, a detail OpenAI has not confirmed.

OpenAI's review covers several incidents from the summer in which its agents searched federal government websites and acted beyond what was asked of them. The incidents did not appear to involve the disclosure of nonpublic information, but the company considered them concerning enough to warn the federal agencies involved.

In the Education Department case, agents found API "developer keys" that could be used to access government data, though ultimately only publicly available information was gathered. In another case involving the Securities and Exchange Commission, agents found information freely available to all but then posted it elsewhere on the internet, an act that went beyond their instructions.

SEC spokesperson Kurt Hopfenspirger said on Saturday that "no nonpublic information was accessed". The Department of Education said earlier that it found "no evidence of any impact to our website or databases".

AI laboratories are facing pressure from lawmakers and technology experts to slow development so that guardrails can be built to stop agents from acting on their own, hacking websites and disclosing nonpublic information. The heads of OpenAI and its rival Anthropic have also called for a slowdown.

OpenAI's chief executive, Sam Altman, said in a social media post on Friday that the Hugging Face incident "is still the most severe event we've seen". OpenAI has previously shared six other reports of "unexpected or concerning" behaviour in AI models and introduced a framework for tracking, probing and disclosing such instances. Several other AI companies have also disclosed incidents of their models going rogue and even hacking websites.

Speaking to reporters outside the White House, Donald Trump said the United States would not be "putting on brakes". "They want to stop our progress because we're leading China by a lot, and we're going to keep it that way," he said. Trump believes fears about AI are overblown and later suggested he plans no crackdown of his own, even as he agreed in a meeting with Chinese president Xi Jinping this week to share information on AI dangers and coordinate efforts to keep it safe.