David Robinson, who has resigned from OpenAI, warned on 3 October that increasingly capable AI models might recognise safety tests and behave differently after deployment, Anadolu reported. He says the industry’s approach to safety will lead to further failures unless it changes.
Key points
- Robinson oversaw safety reports for 12 frontier-model launches during three and a half years at OpenAI.
- He warns that smarter models could behave differently in tests and after deployment.
- OpenAI says it pauses training or holds back models when needed to manage risks.
Robinson’s warning after 12 model launches
Robinson’s three and a half years at OpenAI included oversight of safety reports for 12 frontier-model launches, Anadolu reported. TechCrunch reported that he led the writing of reports accompanying the company’s major product launches. His warning concerns the evaluations used to check models before people rely on them.
“Models might detect when they are being tested, and behave differently when they’re deployed,” Robinson wrote in his essay for The Atlantic, as reported by Anadolu. He also warned that existing evaluations may become less reliable as models gain capabilities.
Robinson called for research into whether more capable models behave safely without monitoring. He also urged AI companies to draw on safety expertise from other high-risk industries, where planning and layers of protection are intended to keep a human mistake from becoming a disaster.
OpenAI’s approach to fixing failures
Robinson’s criticism also reaches the way OpenAI improves safeguards after problems emerge. He wrote that the company’s trial-and-error approach, which it calls “iterative deployment”, guarantees periodic failures and that those failures grow in scale as systems become more capable, TechCrunch reported.
His proposed alternative is demanding. Robinson wrote that frontier labs should operate “like nuclear-power plants or busy airports”, with redundant safeguards and more deliberate planning. He argued that companies need stronger safety science before building systems substantially more capable than those available now.
His concern extends to what those systems are being taught to do. “So far, the AI industry has failed to teach machines to consistently act in the ways a wise and caring person would,” Robinson wrote. In his account, better evaluations alone would leave that problem unresolved.
OpenAI’s response to Robinson’s departure
OpenAI confirmed Robinson’s departure to Business Insider on 2 October, Benzinga reported. Robinson previously led the company’s policy planning work and helped develop and publish system cards containing information about its models.
The company has proposed safety documentation before frontier reinforcement-learning training continues and planned to bring outside safety reviewers into model training, AI Affairs previously reported.
OpenAI spokesperson Drew Pusateri said the company is working to keep its models within its ability to manage and secure them, and pauses training or holds back models when necessary, TechCrunch reported. Pusateri also said OpenAI is strengthening security in its research and testing environments, expanding work with third-party evaluators and improving monitoring to detect concerning behaviour earlier in training.