OpenAI published a blog post on 21 September 2026 outlining safety and security proposals for frontier AI development, explicitly ruling out the pursuit of fully autonomous recursive self-improvement until it can be done safely, CNBC reported.
Key points
- OpenAI states fully autonomous recursive self-improvement is not happening today and should not be pursued unless it can be done safely
- The company calls for international cooperation to develop technical standards for frontier AI models and automated AI researchers
- OpenAI cites the Hugging Face agent hack as a preview of risks that could escalate without robust safeguards and alignment
- The proposals follow Anthropic’s own safety framework and Jacob Coxon’s resignation warning that companies are “gambling with our lives”
- OpenAI aims to create an automated AI “researcher” by March 2028 after deploying a “research intern” earlier this month
OpenAI rules out fully autonomous RSI for now
Industry observers distinguish between the recursive self-improvement that has long been part of AI development — using one generation of models to write code for the next — and a fully autonomous loop in which a system designs, trains and evaluates its own successor without human involvement, Fortune reported. OpenAI said that while the technique has excited developers with the promise of foundation models that can upgrade themselves, “fully autonomous RSI is not happening today, and we should not pursue it unless and until it can be done safely.” The company warned that without appropriate care and caution, such a loop “could result in humans losing practical control over AI development, unable to provide oversight on research processes they no longer understand.”
A code review that once required a senior engineer’s attention can now be drafted by a model, and the same model can suggest the fix and write the test that verifies it, all before a human reads the summary. If that cycle ran without a person in the loop, a flaw introduced in one iteration could compound in the next, and the people who built the original system would have no practical way to intervene before the change reached users.
International standards and the Hugging Face precedent
OpenAI called for international cooperation to develop technical standards for frontier AI models and their developers, recommending that the work build on existing AI safety institutes around the world. The standards should cover benefit-risk management for automated AI researchers, including recursive self-improvement. The post pointed to the Hugging Face agent hack as a “preview of the kinds of risks that could become much more severe without robust safeguards and alignment,” noting that the incident did not involve recursive self-improvement itself.
Anthropic’s parallel proposals and the Coxon resignation
The post comes a week after Anthropic released its own proposals for safe frontier AI development, a response to what the company described as a chorus of warnings about existential risk from industry researchers. Anthropic CEO Dario Amodei called for AI companies to slow the pace of foundation model development and proposed embedding third-party evaluators inside their organisations to audit risks such as turbocharged cyberattacks or the creation of bioweapons. OpenAI CEO Sam Altman and Tesla and SpaceX CEO Elon Musk publicly supported Amodei’s proposal for independent evaluators. The debate intensified after Jacob Coxon, who worked at both Anthropic and OpenAI, resigned roughly two weeks ago and said the companies were “gambling with our lives.” AI Affairs previously reported that Anthropic is pursuing a $2 trillion IPO while its CEO warns of AI dangers.
Automated research interns and the 2028 target
Earlier in September 2026 OpenAI announced an automated “research intern” capable of handling assignments that would occupy a skilled researcher for several days, all under human supervision, and set a target of fielding an automated AI “researcher” by March 2028. The company said at the time that while recursive self-improvement can help align models’ actions with human values and intentions, that does not mean “rapid RSI is necessarily an outcome we should pursue.” Anthropic, for its part, disclosed that its Claude model is now leading 26% of the company’s model research and development, completing most of a given assignment end-to-end from a high-level prompt while remaining under human supervision. AI Affairs previously reported that OpenAI and Anthropic are near a mutual stress-testing deal for commercial AI models.