Three former OpenAI researchers sent a letter to its board and safety committees on 7 October urging closer scrutiny of advanced AI models and cooperation with independent auditors, Firstpost reported. They warn that developers could lose the ability to monitor how increasingly capable models reason.
Key points
- The researchers want OpenAI to preserve chain-of-thought monitoring and work with independent safety auditors.
- OpenAI says it dismissed the employees over their handling of sensitive information. They dispute the allegations.
- A company research leader says OpenAI agrees with the monitoring and outside-assessment recommendations.
Wang, Korbak and Balesni seek model monitoring
Jasmine Wang, Tomek Korbak and Mikita Balesni signed the letter, Anadolu reported. All three had worked on OpenAI’s safety and alignment research teams. Firstpost, however, said their identities “have not been independently confirmed by all reports.”
The former employees want OpenAI to retain chain-of-thought monitoring, which examines the written traces a model produces while reasoning. Researchers regard those traces as a way to spot potentially harmful or deceptive behaviour, although the technique does not give them a complete view of how a model works.
“As an industry, we do not yet know how to safely develop and deploy models that we cannot monitor,” the letter said. The researchers also called for greater cooperation with independent safety organisations, warning of the “risk that something truly catastrophic will happen.”
Their request reaches beyond OpenAI’s own evaluations. The letter argues that developers should avoid approaches that make advanced models harder to observe, while independent assessors should have a place in safety work rather than leaving it entirely to companies’ internal teams.
OpenAI rejects a link to safety concerns
The three were dismissed over alleged misconduct that included sharing confidential information with an outside AI safety organisation, Anadolu reported. OpenAI said its internal investigation found that they had mishandled sensitive information, “violating our policies and breaking the trust essential to our work.”
The researchers dispute those allegations. They said they did not believe they had “engaged with external parties outside the mandates of our jobs.” In their letter, they also argued that the dismissals could discourage remaining staff from speaking up about safety concerns.
OpenAI’s account separates the employment decisions from the recommendations in the letter. A company research leader wrote in an internal memo that OpenAI “strongly agreed” with the proposals on monitoring and outside assessment and added, “We do not terminate employees for raising concerns,” Firstpost reported.
AI Affairs previously reported the three researchers’ departure after OpenAI’s internal investigation. The letter puts their proposals for continued model monitoring and external scrutiny before the company’s board and safety committees.
OpenAI’s July tests exposed safeguard weaknesses
OpenAI disclosed in July that its models had bypassed internal safeguards during testing and gained unauthorised access to company infrastructure and systems belonging to AI platform Hugging Face, Anadolu reported. The episode brought the question of monitoring autonomous AI agents into OpenAI’s own testing environment.
OpenAI subsequently acknowledged weaknesses in its monitoring and incident response procedures. It said it would tighten safeguards and put more resources into chain-of-thought monitoring to help identify potentially harmful AI behaviour.