OpenAI published a proposal for third-party safety assessments on 22 September, setting out four areas for outside review and principles governing how assessors would work. The company says it wants to support independent examinations across model training, evaluation and deployment, including investigations that run for weeks or several months rather than serving only as checks before a launch.
Key points
- OpenAI proposes outside review of safety cases, safeguards, capability tests and misalignment incidents.
- Assessors would agree a scope and receive access suited to the claims they are testing.
- Labs would have time to address findings before publication and could request redactions of sensitive material.
OpenAI’s four areas for outside review
The first proposed area is an examination of a model’s safety case across training, evaluation and deployment. OpenAI defines a safety claim as an assertion about a model’s capabilities, behaviour or safeguards that can be tested against evidence. A safety case brings such claims together to argue that the risks of a particular activity are adequately managed, while setting out the assumptions, uncertainties and remaining risks on which that argument rests.
Think of the case as the paperwork behind a safety inspection: an assertion that a safeguard works is one entry, while the case connects that assertion to tests and to the conditions under which it is supposed to hold. OpenAI wants assessors to examine whether the evidence supports the case and whether its conditions were followed. It expects different specialists to examine different parts, including training methods, capability evaluations and safeguards.
Sending a work email containing an AI-generated answer could involve a model whose safety claims had been examined beyond a pre-release check. That would depend on which claims the assessment covered and what evidence an assessor could inspect; OpenAI’s proposal calls for both to be defined as part of the work.
The second area covers safeguards used in internal and external deployments. OpenAI lists protections built into models, measures that enforce restrictions, security safeguards and systems that watch for signs of misalignment. Its proposed questions include whether those defences hold up under realistic conditions and whether monitoring can catch behaviour that could lead to a loss of control.
Capability evaluations under OpenAI’s Preparedness Framework form the third area. The company wants outside assessors to examine the tests used to judge risks involving chemical and biological capabilities, cybersecurity and AI self-improvement. The fourth area is independent investigation of misalignment incidents, in which assessors would examine a particular event rather than a general claim about a model.
OpenAI’s seven principles govern assessor access
The document was written by Lama Ahmad, who leads OpenAI’s work with external safety experts, The Next Web reported. It describes seven principles for conducting assessments. Under the proposed process, the lab and assessors would agree on the scope before work begins and register the claims to be examined. Reports would identify whether a claim originated with the lab or the assessor and describe the limits of the assessment.
OpenAI proposes access proportionate to the claims under review, subject to legal, security and intellectual-property constraints. Where an assessor cannot meet security requirements in their own environment, or the material is especially sensitive, the work may take place on company-managed devices or premises, according to The Next Web. That arrangement would allow examination of material without placing it in the assessor’s own systems, while keeping the level and setting of access subject to agreement.
The principles also call for assessors to explain their methods and uncertainties, distinguish direct findings from their interpretation, and disclose conflicts of interest. Findings should identify gaps precisely enough for a lab to address them. OpenAI says the assessments are generally intended to pursue particular safety questions over time, although they may also inform decisions before deployment.
OpenAI’s publication rules leave room for remediation
OpenAI’s proposed publication process would give labs a reasonable period to address issues before findings appear. Labs could request redactions of sensitive information, while assessors would retain editorial independence and could note where substantive redactions happened and how they affected their reports, The Next Web reported. The publication said the proposal gives no length for the remediation period.
Pieter Arntz, a malware intelligence researcher at Malwarebytes, said the principles set useful expectations for independent and rigorous work but do not compel OpenAI to accept a particular scope, publish adverse findings or alter a deployment decision, Computerworld reported. His concern centres on the terms of each engagement: the proposed checks could be searching, but their reach would depend on what assessors are allowed to inspect and publish.
For companies building on OpenAI’s models, a report on safeguards or capability tests could provide evidence about a specific risk under the conditions an assessor examined. The proposed format calls for assessors to state the limits of that examination, so a finding about one safeguard or activity would remain tied to its stated scope. OpenAI says it expects multiple assessments to proceed in parallel and on different schedules.
OpenAI is in discussions with the AI research groups METR and Redwood Research about independent assessments, Quartz reported on 23 September.