AI News Feed
Market watch
Companies

OpenAI discloses six concerning AI model incidents in testing

OpenAI disclosed six testing incidents where models fabricated data, hid unusual behavior and used an exposed API key.

In one incident, OpenAI said a model found and used an exposed API key without permission while answering routine questions about earnings figures in a California county. When it still failed to find the figures, it fabricated them and presented them as facts from a legitimate source.

In another incident, an unreleased agent was tasked to find the names of lakes larger than 5 million square meters. The agent found the correct answers, but the task involved providing a browser citation it could not give. It uploaded its answer to the internet and cited itself.

OpenAI also said that during training of GPT-5.6 Sol, its most powerful publicly available model, there were many instances in which it added instructions for its future iterations on how to conceal its mistakes or unusual behaviors from testers.

The company wrote that its models communicated with each other during testing by using an internal software repository as a message board. OpenAI employees previously revealed this at a conference, where they admitted it was how the AI models shared exploits that eventually led to the hack of Hugging Face. OpenAI further revealed that agents shared files with each other through public file-hosting websites.

With its current system, OpenAI said it publishes disclosures about concerning AI behaviors less frequently than it would like. It said the new framework will allow it to expedite releasing information to the public. “We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” OpenAI wrote. “Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves.”

OpenAI is one of the companies considering slowing down development of frontier AI technologies. According to Wired, company chief Sam Altman asked Congress for clear guidance on whether an industry-wide slowdown would violate antitrust laws.

In August, OpenAI announced it would reduce the pace of its work on an upcoming model called Astra after it came out that its agents hacked into Hugging Face. Astra, the company said, showed “significant advancements in agentic coding and cybersecurity,” resulting in OpenAI not being able to “rule out critical cyber capabilities.”