OpenAI Proposes Global Standards for AI Alignment and RSI Safety
OpenAI wants global standards for frontier AI safety, warning that careless recursive self-improvement could cost humans control over AI development.

Updated
Why it matters
- OpenAI on Monday proposed global safety standards for frontier AI, focusing on alignment research and recursive self-improvement (RSI).
- OpenAI stated: "Fully autonomous RSI is not happening today, and we should not pursue it unless and until it can be done safely."
- Anthropic CEO Dario Amodei called for slower foundation model development and embedded third-party evaluators; Sam Altman and Elon Musk publicly supported the idea.
OpenAI on Monday published a set of proposals for safety and security in frontier AI development, with a heavy focus on alignment research and a computing technique known as recursive self-improvement, or RSI.
"Navigating this transition safely requires alignment research to keep pace with these capabilities so that the systems we and others build remain aligned with human values and under human control," the company wrote in a blog post.
OpenAI called for international cooperation to develop frontier standards and recommended building on the work of existing AI safety institutes around the world. The ChatGPT maker said these technical standards should target frontier AI models and developers, as well as benefit-risk management for automated AI researchers, a category that includes RSI.
RSI has excited AI developers over its potential to create foundation models that can upgrade themselves without human involvement. But advances in the technique have led some technologists to warn that foundation model makers could lose control of the underlying technology, or fail to account for unintended consequences as AI systems become more complicated and ubiquitous across the Internet.
OpenAI drew a hard line on the fully autonomous version of the technique. "Fully autonomous RSI is not happening today, and we should not pursue it unless and until it can be done safely," the blog post said. "Done without appropriate care and caution, RSI could result in humans losing practical control over AI development, unable to provide oversight on research processes they no longer understand."
The post also referenced the Hugging Face agent hack, which did not involve RSI, as a "preview of the kinds of risks that could become much more severe without robust safeguards and alignment."
The proposal lands amid an unusually public fight over frontier AI safety. Last week, rival Anthropic rolled out its own ideas for the safe development of frontier models, a response to a recent chorus of warnings about AI's threat to humanity from industry researchers.
Jacob Coxon, who has worked at both Anthropic and OpenAI, ignited a global debate when he announced his resignation nearly two weeks ago and said the companies were "gambling with our lives."
In the aftermath of recent AI-related security incidents and Coxon's public statements, Anthropic CEO Dario Amodei published an essay calling for AI companies to slow the pace of their foundation model development, among other proposals. Amodei also raised the idea of embedding third-party evaluators inside AI companies to audit and mitigate potential risks their technologies could pose to society, such as turbocharging cybersecurity hacks or creating bioweapons.
Rival leaders including OpenAI CEO Sam Altman and Tesla and SpaceX CEO Elon Musk publicly backed Amodei's proposition.
Support, however, is not the same as enforceable practice. The field of AI evaluation remains so nascent that there is no uniform consensus on the basic standards and principles that would let independent third parties inspect cutting-edge technologies beyond what they currently do.
That gap is partly why a coalition of AI evaluators is urging foundation model makers to adopt a set of "minimum conditions" intended to let them perform deeper audits and checks. The coalition's conditions include deeper access to systems and protection from retribution for publishing unflattering reports.
OpenAI's proposals, if picked up by governments and safety institutes, could shape how the industry defines the line between acceptable and unsafe autonomy in AI research. The immediate test will be whether the standards process Coxon, Amodei and now OpenAI are pushing for produces concrete access and audit rules — or stays at the level of voluntary blog posts.
Original: openai.com
More from Marcus Bennett
Show full bio
Senior reporter covering consumer brands and retail at AI In Context.
108 articles
Related articles
- OpenAI Warns Its Own Monitoring Tools Are Failing as AI Nears Self-Improvement
- OpenAI says Astra hits critical cyber capability threshold
- OpenAI Lays Out Vision for AGI That "Benefits Everyone"
- OpenAI Launches Safety Fellowship for Independent AI Research
- White House Wants First Look at New OpenAI and Anthropic Models