OpenAI Proposes Global AI Safety Standards for Alignment, RSI
OpenAI proposed global frontier AI safety standards focused on alignment and recursive self-improvement, urging international cooperation and caution on autonomous RSI.
The ChatGPT maker said in a blog post that “navigating this transition safely requires alignment research to keep pace with these capabilities so that the systems we and others build remain aligned with human values and under human control.” OpenAI recommended building on the work of existing AI safety institutes around the world.
OpenAI said the technical standards should focus on frontier AI models and developers, as well as benefit-risk management for automated AI researchers, which includes RSI. RSI has excited AI developers over its potential to create foundation models that can upgrade themselves without human involvement.
Advancements within RSI have led some technologists to raise concerns that foundation model makers could lose control of the underlying technology or fail to account for potential unintended consequences as the AI systems become more complicated and ubiquitous across the Internet. OpenAI addressed that concern directly. “Fully autonomous RSI is not happening today, and we should not pursue it unless and until it can be done safely,” the blog post said. “Done without appropriate care and caution, RSI could result in humans losing practical control over AI development, unable to provide oversight on research processes they no longer understand.”
The proposal comes amid continuing debate over frontier AI risks. Last week, rival Anthropic rolled out its own ideas for the safe development of frontier AI models. That followed a chorus of warnings from industry researchers about AI’s threat to humanity. Jacob Coxon, who has worked at both Anthropic and OpenAI, ignited a global debate when he announced his resignation nearly two weeks ago and said the companies were “gambling with our lives.”
Editor's Summary
OpenAI proposed international standards for frontier AI safety, with alignment and RSI as central concerns. The company said fully autonomous RSI should not be pursued unless it can be done safely. The proposal follows Anthropic’s own safety ideas and a former employee’s public resignation that intensified debate over frontier AI risks.