AI Affairs, home

Friday 25 September 2026

Technology

OpenAI plans to bring outside safety reviewers into model training

The company proposes longer-running reviews of safety arguments, safeguards and risky capabilities, with outside groups potentially examining models well before a launch decision.

Sam Altman and Masayoshi Son meeting at Japanese Prime Minister's Residence
Photo: 首相官邸ホームページ / Office of the Prime Minister of Japan Official Website, CC BY 4.0, via Wikimedia Commons (cropped)

OpenAI said on 22 September 2026 that it would support independent technical safety assessments across the training, evaluation and deployment of its AI models, giving outside organisations access earlier in development than its pre-launch reviews typically allowed. The company set out four priorities for deeper scrutiny, including reviews of the evidence behind its safety claims and checks on whether safeguards work in practice.

Key points

  • OpenAI proposes outside assessments during training, evaluation and deployment, rather than concentrating them near launch.
  • Its priorities include safety cases, deployed safeguards, evaluations of dangerous capabilities and investigations of serious misalignment incidents.
  • OpenAI is discussing the work with METR and Redwood Research, Quartz reported.

OpenAI shifts scrutiny into training

OpenAI has previously given outside assessors access at various points in development and deployment, including information about technical safeguards and confidential material for incident response. Lama Ahmad, who oversees its relationships with external safety experts, said third-party reviews had nevertheless been concentrated on safety checks and capability evaluations before launch. The new proposal would let assessors examine questions arising while a model is being trained, as well as when it is evaluated or used.

That timing changes the work an assessor could undertake. A review near release can test the model and safeguards available at that point. OpenAI says longer-running access would allow outside groups to challenge its assumptions, identify risks it may have missed and reach their own conclusions about its safeguards across development. The company expects several assessments to proceed at once, with some lasting weeks and others several months.

OpenAI describes the proposed work as generally independent of a particular launch. It says third-party assessments may also form part of pre-deployment checks and inform deployment decisions, but the longer engagements would examine specific safety claims over time. That distinction matters to companies building on a model: a launch check addresses what is about to be released, while the proposed reviews could also examine the training and internal use behind it.

Safety cases put OpenAI’s evidence under review

OpenAI’s first priority is independent assessment of its safety cases across training, evaluation and both internal and external deployment. A safety case is the company’s organised argument that particular risks are adequately controlled for a particular activity. Think of an inspection file: it links each assertion about what is safe to the evidence for it, while recording the conditions and remaining risks that could affect the conclusion. OpenAI says assessors could examine the whole case or separate parts requiring different expertise.

Checking a work answer before passing it on could depend on whether the safeguards around it hold up where the answer is produced. An assessment extending into deployment could examine those conditions alongside earlier tests, if OpenAI carries out the reviews it proposes.

The company wants assessors to ask whether evidence backs a safety case and whether its stated conditions were followed during training and deployment. It also proposes examining training methods for incentives that might reward deception, attempts to exploit a scoring system or efforts to evade restrictions. This is a deeper question than whether a model passes a single test: the assessor would have to look at the claim, the supporting evidence and the circumstances in which that evidence applies.

A second priority covers critical safeguards in internal and external deployments. OpenAI says those protections include measures built into models, enforcement and security controls, and monitors for misaligned behaviour. Its remaining priorities call for outsiders to scrutinise tests of AI self-improvement, cybersecurity and chemical and biological dangers, and to investigate critical misalignment incidents independently, Quartz reported.

METR and Redwood Research discuss access

OpenAI is in discussions with METR and Redwood Research, among other AI research groups, about the assessments. For particularly sensitive work, outside evaluators may be brought into OpenAI’s offices. Sam Altman said on 12 September that independent evaluators would have access to OpenAI’s offices, receive equipment and be allowed to publish what they found, Quartz reported.

OpenAI says effective assessments require independence, scientific rigour, strong security practices and clear responsibilities. Access to training and internal deployment could expose assessors to confidential information, and the company says laboratories must enable scrutiny while protecting sensitive material. Its proposal calls for shared safety and security standards between laboratories and assessors; it also distinguishes these private and non-profit technical reviews from its work with governments on testing.

The 18 September letter asks for protections

The question of how much access outsiders need extends beyond OpenAI. Anthropic chief executive Dario Amodei proposed giving evaluators employee-like access to inspect advanced models and development processes. More than 100 AI experts and evaluators signed a public letter on 18 September calling for scientific objectivity, transparency, independence and protections for those doing the work, CNBC reported.

Anthropic has said it would embed Accenture evaluators to test its models. OpenAI and Anthropic had also neared a mutual model stress-testing deal, AI Affairs previously reported. Embedded assessment asks more of an outside group than testing a finished model: the access OpenAI proposes would extend to evidence about how risks are identified and managed during development.

Conrad Stosz, chair of the AI Evaluator Forum consortium that organised the letter, said outside assessments should not replace a developer’s own testing or efforts to reduce risks. The signatories also want evaluators shielded from retaliation by the companies they work with, CNBC reported.

Topics: Foundation models, Safety