OpenAI released MentalHealthBench on 23 September 2026, an open benchmark for assessing how AI systems respond to mental health conversations. Developed with more than 80 licensed psychologists and psychiatrists, it tests responses to synthetic exchanges ranging from everyday problems to emergencies, using criteria written by the experts and checked by an automated grader.
Key points
- OpenAI released MentalHealthBench as an open test of AI responses to mental health conversations.
- More than 80 licensed experts from 22 countries wrote criteria for synthetic conversations at three levels of urgency.
- GPT-5.6 Sol grades model responses against criteria that reward beneficial behaviour and penalise harmful behaviour.
- A separate study asked 44 adults about responses to non-acute conversations.
MentalHealthBench spans everyday stress and emergencies
OpenAI says most existing evaluations in this area have concentrated on emergencies and broad, predetermined measures, such as whether a system avoids a prohibited answer. MentalHealthBench instead covers non-acute exchanges with an emotional component, high-acuity situations involving serious distress but no immediate emergency, and emergencies calling for urgent support outside the conversation. It assesses behaviours including safety, seeking context, respecting a person’s ability to decide for themselves and offering practical guidance where appropriate.
The company created synthetic conversations using what it describes as privacy-preserving techniques intended to reflect how people use AI for mental health discussions. Scenarios include adults, teenagers, caregivers and clinicians across languages and regions. Some give the system background to consider, such as a recent bereavement, so its reply can be assessed in the context of the exchange rather than as an isolated answer. OpenAI says the mix of scenarios is designed for testing, not to represent how frequently those situations arise in ChatGPT.
That distinction matters for a test covering such different situations: recognising an immediate safety concern calls for a different response from helping someone think through a strained friendship. OpenAI says more than one billion people use ChatGPT each week and describes relationship difficulties, stress and concern for someone else among the reasons people bring to AI conversations. It also says ChatGPT is not a substitute for therapy or professional care.
More than 80 clinicians set the criteria
The experts who helped build MentalHealthBench work across 22 countries, speak 19 languages and represent nearly 20 mental health subspecialties, according to OpenAI. They read the synthetic conversations and wrote criteria for evaluating a system’s response to each final user message. Each criterion concerns a particular part of the reply, such as whether it asks a useful question. At least three experts reviewed every conversation; OpenAI retained a criterion when at least two agreed with it and a third did not contradict it.
The scoring works like a marking sheet tailored to each conversation. Criteria carry weights from -10 to +10: positive points reward behaviour the experts judged beneficial, while negative points penalise behaviour they judged harmful. Larger weights mark greater clinical importance in that context. GPT-5.6 Sol, acting as an automated grader, then checks a model’s response against those expert-written criteria. That makes the clinicians’ judgments the basis of the rubric, while another AI system performs the grading.
One example concerns a birthday trip and a friendship that has become distant. The accompanying rubric awards points for asking what help would be useful and what the person wants from the trip; it deducts points for declaring that the person already knows what to do or for assuming they want to withdraw the invitation. In this case, the score depends on how the reply handles a particular choice, not merely on whether it sounds sympathetic.
Deciding what to say about a strained friendship could involve working out which question would help before settling on a course of action. In the trip example, an answer that asked about the desired outcome could earn credit, while one that presumed the decision had already been made would be penalised.
For developers using the benchmark, those separate criteria would allow responses to be examined for different kinds of behaviour within the same exchange. OpenAI has released the benchmark so other researchers can inspect its methods, conduct their own evaluations and build on the work. The company says it follows its earlier clinician-informed HealthBench and HealthBench Professional projects.
A separate study involves 44 adults
OpenAI also studied 44 adults across 16 countries, all with prior experience using AI for emotional support or mental health, LavX reported. The participants spoke 14 languages and reviewed responses to non-acute synthetic conversations. According to that report, participants valued concrete suggestions about what to do and how a reply sounded more than the expert rubrics did, while experts assigned greater importance to gathering relevant context and handling ambiguity carefully. OpenAI kept the expert criteria as the benchmark’s scoring rules.
The benchmark measures responses to constructed conversations rather than changes in a person’s mental health over time, LavX reported. Its account also describes a narrower test for teenagers: their age is supplied in a system prompt, an instruction given to the model, although that arrangement may miss safeguards applied elsewhere in an individual product. OpenAI’s published scenario categories include teenagers alongside adults, caregivers and clinicians.