OpenAI Safety Transparency Lead Departs as Three Safety Employees Are Fired
David Robinson, who led safety transparency in OpenAI's Safety Systems team, has left the company, and The Wall Street Journal reported that three safety and alignment researchers were fired for sharing sensitive information with an outside AI safety organization.
Robinson led work intended to help the public understand and trust OpenAI's technical safety efforts. His team acted as a translator between technical departments and the public, turning internal technical findings, risk judgments and deployment decisions into public materials. The transparency function sat within the Trustworthy AI team inside Safety Systems, and its output included system cards, the Deployment Safety Hub, safety research blogs and public governance documents.
The most visible product was the system card. According to the report, these documents describe how a model was trained, what safety evaluations it underwent, what capabilities it reached in high-risk areas such as biology and cyberattacks, what internal safeguards were applied, and what risks remained after those safeguards. The report described the cards as dozens of pages long on average and compared them to a nutrition label on food packaging: few people may read them, but they are required. Robinson said the work was OpenAI's most detailed and comprehensive public explanation of its technical safety work.
Robinson was also a principal drafter of OpenAI's Preparedness Framework 2.0. The framework tracks severe risks from frontier models and sets out in advance what evaluations and protections are required when a model reaches certain capabilities. The second version added three categories of capabilities to track: biological and chemical capabilities, cybersecurity capabilities, and AI self-improvement capabilities. If tracking finds a model has reached High capability, it may significantly amplify existing severe harm pathways and must have sufficient safeguards before deployment. If it reaches Critical, it may create severe harm pathways from scratch, raising questions about whether development should continue.
The report cited Deep Research as a case in point. In the Deep Research System Card, the model's biological evaluation showed it was approaching the High capability threshold, meaning it was nearing the level at which it could materially help non-experts create known biological threats. Describing such threshold cases accurately to the public was part of Robinson's responsibility.
Shortly before his departure, Robinson was recruiting OpenAI's first Safety Transparency Editor. The role was to oversee the editorial quality of important transparency materials and push the narrative and presentation of system cards and related documents from project launch to final publication. The position was still listed as open at the time of the report, and it was not known whether anyone had been hired. Robinson has not publicly explained his departure, and OpenAI has not announced a successor.
Before news of Robinson's departure emerged, OpenAI had already fired three safety-focused employees. The Wall Street Journal reported that they were Jasmine Wang, Tomek Korbak and Mikita Balesni, who mainly worked on safety and model alignment research. OpenAI said an internal investigation found that the three shared sensitive company information with an external AI safety organization, including content about OpenAI infrastructure architecture. Tomek Korbak had previously served as OpenAI's technical contact with third-party evaluators METR and Redwood Research. Those organizations were invited into OpenAI to investigate after the Hugging Face incident, and Korbak was the person who liaised with their researchers.
OpenAI's safety team has seen frequent turnover since the Dario Amodei era. Recent examples cited in the report include OpenAI disbanding its catastrophic risk emergency team and GPU expert Scott Gray leaving, four senior executives departing in one month including a former COO, and a safety head leaving after the release of GPT-5.6. The report said the number of departures was too large to list individually.
Before leaving, Robinson spoke publicly on X about the risks of highly capable AI. In early September, he said people inside OpenAI were gradually realizing the impact highly capable AI models could have. He said OpenAI was indeed undergoing major changes every day, but he did not know whether the changes were fast enough. In his latest X post, he reposted another researcher and said he agreed completely: he did not think racing toward RSI was a prisoner's dilemma, and it might be as crazy and arrogant as ordinary people intuitively see it. Robinson told AI staff that if they chose to stay, they should think about how to use their influence inside the company to push it to address AI risk more seriously.