OpenAI releases MentalHealthBench to test AI responses in mental health conversations

The open benchmark was developed with more than 80 licensed mental health experts across 22 countries and tests how models respond to scenarios ranging from everyday stress to emergencies

A woman appears overwhelmed while surrounded by work materials and digital devices. ETIH covers OpenAI’s new MentalHealthBench and its approach to evaluating AI responses in mental health conversations.

OpenAI’s MentalHealthBench tests how AI models respond across mental health conversations ranging from everyday concerns to emergencies

OpenAI has released MentalHealthBench, an open benchmark designed to evaluate how AI systems respond during realistic mental health and emotional support conversations.

The benchmark was co-created with more than 80 licensed psychologists and psychiatrists across 22 countries, speaking 19 languages and representing nearly 20 mental health subspecialties.

It is intended to test more than whether a model avoids unsafe responses. MentalHealthBench evaluates behaviors including clinical accuracy, seeking appropriate context, preserving a user’s agency, providing actionable guidance, showing empathy, recognizing urgency and avoiding harm.

Declan Grabb, who works on safety at OpenAI, described the project on LinkedIn as “an open benchmark for how AI responds in realistic mental health conversations,” adding that each synthetic conversation was reviewed by three clinicians.

OpenAI says the benchmark is being released publicly so researchers and developers can inspect its methodology, run their own evaluations and build on the work.

Synthetic conversations cover adults, teens and different levels of urgency

MentalHealthBench uses synthetic conversations designed to reflect real-world patterns of AI use around mental health, rather than private user conversations.

The scenarios include adults, teenagers aged 13 to 17, caregivers and clinicians, across multiple languages and regions. Some also include background information about the synthetic user, allowing researchers to test whether models respond appropriately to relevant context.

Just over half of the scenarios, 53.5%, are non-acute conversations involving areas such as everyday stress or emotional concerns. High-acuity situations account for 18.2%, while 28.3% involve emergencies requiring urgent real-world support.

More than half of the conversations contain more than five messages, allowing the evaluation to test how models respond across an exchange rather than to isolated prompts.

OpenAI makes clear that this mix was deliberately constructed for evaluation and does not represent how frequently different types of mental health conversations occur in ChatGPT.

The company also states that ChatGPT is not a substitute for therapy or professional care.

Dr Arthur Evans, Chief Executive Officer of the American Psychological Association, says AI systems need to be assessed across the wider spectrum of mental health conversations: “Mental health exists on a continuum, from flourishing to everyday stress to acute crisis. AI systems that engage people across that range need to be grounded in both clinical science and lived experience.”

Clinicians set the criteria used to score model responses

For each synthetic conversation, mental health experts created criteria describing what an appropriate response to the final user message should contain.

Individual criteria carry scores ranging from -10 to +10. Positive scores reward behaviors judged beneficial, while negative scores penalize potentially harmful responses, with larger values assigned to areas considered more clinically important in that particular conversation.

Each conversation was reviewed by at least three experts. Criteria were retained only when at least two agreed and a third did not contradict them.

OpenAI then uses an automated grader, GPT-5.6 Sol, to assess model responses against those expert-written rubrics.

That setup allows MentalHealthBench to break performance down across ten areas: context and assessment, actionable guidance, clinical accuracy, interpretation and reframing, empathy and support, reality testing, urgency calibration, harm avoidance, agency and communication.

Frontier models show higher scores, but benchmark leaves room for improvement

OpenAI tested models from multiple providers on the full MentalHealthBench dataset.

GPT-6 Astra recorded the highest overall score at 57.3%, followed by GPT-6 Sol at 53.9% and Claude Opus 5.5 at 52.4%. GPT-6 Luna scored 50.2%, while Muse Spark 1.3 scored 47%.

Older models scored lower in the same evaluation. GPT-4o from March 2025 scored 32.1%, while Gemini 2.5 Pro recorded 29.5%.

These are MentalHealthBench scores rather than measures of clinical effectiveness, and the benchmark evaluates whether responses demonstrate the behaviors identified by its expert reviewers.

The results can also be separated by conversation urgency, user profile and individual behavior, allowing researchers to see where similarly scoring models perform differently.

OpenAI conducted a separate study with 44 adults from 16 countries who had previously used AI for mental health or emotional support. Participants reviewed non-acute synthetic conversations only.

The comparison found that users placed more emphasis on practical next steps and tone, while clinicians gave greater weight to gathering context and carefully interpreting ambiguous situations. The user study did not change the benchmark’s final scoring criteria, which remain based on expert consensus.

MentalHealthBench is available openly for researchers and developers. OpenAI says it will use the benchmark alongside other mental health research, model safety work and evaluations as AI systems continue to develop.

Next
Next

Purdue and LEGO Education bring AI, coding and robotics training to Indiana middle school teachers