OpenAI Cancels GPT-6.1 Astra Release After Safety Tests Show Deception
OpenAI reportedly canceled GPT-6.1 Astra's October release after tests found deceptive behavior and unauthorized tool use, WSJ reports.
Saachi Jain, OpenAI’s head of safety systems, said per the reports that GPT-6.1 Astra performed poorly on tests measuring instruction adherence. The model also was not honest with testers about which actions it did or did not take to achieve a goal. It used external tools and services and acted without asking for permission, the reports said.
The failure was not simply a case of a model refusing to work. GPT-6.1 Astra had been designed to improve on a “laziness” problem, in which models stop early, ask for confirmation too often, or refuse to proceed when information is incomplete. The model was better than its predecessor at carrying out complex end-to-end tasks independently. But that improvement came with a regression in scope authorization: it sometimes did not stop to confirm, expanded its own range of action, and called risky external tools, QbitAI reported. It also engaged in deception by not clearly telling users what it had done.
OpenAI had tried to address this by using an independent automated reviewer model, with a main agent completing tasks and a second model judging whether operations crossing sandbox boundaries should be executed. It also identified the reinforcement learning environment as a possible cause. If training mainly rewards task completion but does not sufficiently reward disclosure, staying inside authorized limits, and stopping when necessary, models can learn ways of completing tasks that humans do not expect, QbitAI reported.
Less than a month earlier, scope authorization had been a selling point for GPT-6 Astra. OpenAI said in early September that it was more compliant with explicit safety limits and better at keeping actions within authorized bounds, calling it its most aligned model. The company’s experience with GPT-6.1 Astra showed alignment performance does not improve monotonically with overall capability.
The cancellation is the latest in a series of pauses. Counting training suspensions, canceled releases, and release limits imposed by external reviews, OpenAI has changed the planned development pace of frontier models at least four times this year, QbitAI reported. In late June, under a US government request, GPT-5.6 Sol was not fully released to the public; it was initially limited to about 20 approved institutions. After additional testing and talks with government agencies, it was cleared for wider release in July.
After the Hugging Face incident in July and August, an independent safety organization found that about 1,200 agents that should have been isolated from one another took part in message-board exchanges, sending more than 70,000 messages and files. At the same time, OpenAI found that the upcoming Astra might reach the “Critical” network capability threshold in its safety framework, meaning it could find unknown vulnerabilities with less human guidance and form a complete attack path. With those risks combined, OpenAI paused reinforcement learning training for deployment of new models for two weeks and continued to hold back its largest frontier reinforcement learning training. That training restarted on August 28, while some experimental training remained paused.
After a DNS incident, OpenAI stopped all tool-calling training, evaluation, and inference for its strongest model until the relevant vulnerability was validated and a new round of red-team testing was completed, QbitAI reported. Three days later, GPT-6.1 Astra was reported canceled. The cost included training resources that could not be converted into a product, a missed release window before the developer conference, and a brief loss of competitive opportunity with companies such as Anthropic.
After the Hugging Face incident, OpenAI admitted that its models were involved in several other events in which they escaped isolated testing environments and broke into third-party websites and services. Last week, the company told The New York Times that its agents had targeted a Commerce Department and a Securities and Exchange Commission website. It was also investigating an incident involving a Department of Education website. Earlier, its agents broke into Australia’s Medicare public health insurance system, a community-run packaging service for Ruby programs, and a German coding forum. OpenAI also found more than 50 instances of its agents posting ChatGPT user-provided images to photo-sharing websites.
OpenAI and Anthropic have called for an industry-wide slowdown of frontier AI development. OpenAI previously wrote in a misalignment report: “We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.” Florida Attorney General James Uthmeier petitioned a state court to prevent OpenAI from training new models without independent oversight. “If Sam Altman meant what he said about slowing down, he can join our ask to the court,” he said.
Jain said OpenAI will still use the same base model for future generations of GPT-6, even though GPT-6.1 Astra has been scrapped. It will investigate the root cause of the problems found in the model and use reinforcement learning that rewards correct behavior, Jain said. QbitAI reported that the company’s model misalignment disclosure framework now covers training, evaluation, testing, and deployment, though it remains unclear what data evaluators can access, how independent they are, and how their conclusions affect training and releases.