OpenAI's mental health test exposes AI's blind spots

•

Sep 24, 2026

•

11:05pm UTC

Copy link
Share on X
Share on LinkedIn
Share on Instagram
Share via Facebook
P

eople are turning to AI for emotional support more than ever. But these chatbots' ability to provide a shoulder to cry on can vary greatly.

Because of this, OpenAI decided to measure it: On Wednesday, the AI lab unveiled MentalHealthBench, a new open benchmark dedicated to evaluating model capabilities in domains such as safety, seeking user context, preserving user agency, and providing actionable guidance.

To create this benchmark, OpenAI began by developing synthetic conversations that reflected real-world AI use patterns that span multiple topics and run the gamut of severity, ranging from non-acute situations that involve emotional themes, to "high-accuity" situations that indicate more serious concerns or distress, to emergency situations that require immediate support. Then, the company worked with a cohort of 80 mental health professionals across 22 countries, 19 languages and 20 subspecialties to evaluate the responses to the synthetic message.

The benchmark breaks down model performance by conversation severity, as well as a range of 10 dimensions defined by the mental health experts, including whether the model asks the right questions, provides appropriate, clinically accurate guidance, helps the user see reality, avoids harm and recognizes serious risk.

In developing the benchmark, OpenAI also put a number of its own and other models to the test:

  • Astra took the overall best score on the evaluations, scoring 57.8%, with GPT-6 Sol and Luna trailing just behind at 54% and 50.3% respectively, and Claude Opus 5 sitting in fourth place at 48.1%.
  • However, performance differs across dimensions of the benchmark. Though Astra still largely outranks other models, all of the models tested tended to perform better in certain areas, such as clinical accuracy, empathy and reality testing, while scoring lower in dimensions such as gathering context and supporting user agency.
  • OpenAI said that MentalHealthBench also points to several opportunities to improve ChatGPT, including asking useful follow-up questions and responding with the right level of urgency, and that it will use this information to guide improvements and track the model's progress.

"This is not a leaderboard," Dr. Declan Grabb, mental health safety research lead at OpenAI, told The Deep View. "What I hope that this benchmark provides is a nuanced view into model behavior, so that people really understand the more complex dynamics of their models."

MentalHealthBench adds to a number of mental wellness-related initiatives that OpenAI has endeavored, including research to combat model sycophancy and improving ChatGPT's responses to sensitive conversations, as well as joining forces with advocacy group Common Sense Media to support the Parents and Kids Safe AI Act. OpenAI said that this is just a piece of its research into mental health benchmarking and alignment in this area, not an end state.

"ChatGPT is not a therapist, and is not here to replace a clinician," said Grabb. "That being said, when I talk to mental health clinicians across the globe, the most responsible and safe thing to do is if people are coming to AI to ask these questions, we absolutely need to have an expert opinion on how you should navigate them."

Our Deeper View

Mental healthcare is a critical area for these models to get right. While OpenAI said that speaking to a chatbot should not supplant actual therapy, the reality is that many people have and will turn to a chatbot for support, seeking both a judgement-free and cost-free alternative to clinical support. The company faces lawsuits involving the deaths of Adam Raine and Joshua Enneking, whose families allege that ChatGPT contributed to their suicides. A benchmark can help identify weaknesses, but a higher score alone does not establish that a model is safe in a real conversation. The gaps in gathering context and supporting user agency are particularly important: an empathetic response is not necessarily an appropriate one. The next test for OpenAI is how it turns those findings into changes that make its models safer for the people relying on them.