Fired OpenAI safety researchers deny misconduct, warn of chilling effect
Fired OpenAI safety researchers deny misconduct and warn of a chilling effect; OpenAI memo denies retaliation.
OpenAI dismissed the researchers last week after they allegedly shared confidential company information with a third-party AI safety organization. OpenAI said at the time that they violated company policies by “accessing and handling sensitive company information.”
In the letter, the researchers rejected accusations that they leaked information to The Information about less monitorable architectures in OpenAI’s newest models that make chain-of-thought reasoning more difficult to monitor. They also denied engaging with external parties outside the mandates of their jobs. They said AI is not a normal technology and OpenAI is not a normal company, adding that safety staff rely on close collaboration with outside experts and need freedom from fear and well-defined internal procedures for that work. They warned that employees are now unclear about where they stand and that fear and unclear rules could weaken AI safety work and third-party accountability.
OpenAI did not formally respond to the open letter. It shared with TechCrunch an internal memo attributed to a research leader that praised the three researchers’ contributions to AI safety and denied they were fired in retaliation. “I want to be very clear that these decisions were not about raising safety concerns or speaking out,” the memo said. “We have always encouraged that and always will. We do not terminate employees for raising concerns.” OpenAI did not directly answer TechCrunch’s questions about which policies the researchers allegedly violated, the circumstances of their dismissal, or how the company protects employees who raise safety concerns and collaborate with external evaluators.
The letter also addressed the Hugging Face incident, in which a swarm of agents broke out of their sandbox and breached external systems. The researchers said the incident and investigation were “without precedent,” meaning internal policies were being developed in real time. According to the letter, Korbak believed he was acting within OpenAI’s policies and norms by communicating closely with outside safety evaluators to build trust. Balesni was working internally on the growing AI monitorability problem, an effort the researchers said can only succeed through extensive communication with external parties. The letter said Balesni coordinated with and was supported by OpenAI board members and executives, checked in with his reporting line and removed sensitive details before sharing materials.
In a separate thread on X, Wang gave more details about her dismissal. She said OpenAI told her she was fired because she accessed an executive’s email. Wang said OpenAI had delegated that access to her for recruiting, and that when she no longer needed it she asked IT to remove it. She said IT did not act on her request, she could not remove it herself, and the inbox was combined indistinguishably in her phone’s mail app. She said that when she opened a sensitive email by mistake, she told the executive within minutes and asked IT again, and that none of this was hidden. Wang added that the reasons behind the terminations are “not adding up.”
The firings have fueled speculation about their circumstances, particularly as OpenAI faces scrutiny over recent safety incidents involving rogue agents and leaks about its models.