It is relatively easy to test whether a model refused a dangerous answer. It is much harder to tell whether it asked the right question, understood an ambiguous situation or left room for a person to make their own decision. OpenAI wants to make that less mechanical part of mental-health conversations measurable with MentalHealthBench, published on September 23.
The open benchmark was built with more than 80 licensed psychologists and psychiatrists from 22 countries, working across 19 languages. Rather than focusing only on emergencies, it covers everyday conversations, higher-acuity situations and crises. Its scenarios represent adults, teenagers, caregivers and clinicians.
That scope matters because people do not turn to AI only in extreme moments. They also ask how to set a boundary, support someone close to them or prepare for a difficult conversation. In those cases, an answer can be factually reasonable and still feel rushed, generic or intrusive.
The rubric changes with the conversation
The most interesting part of MentalHealthBench is that it does not rely on one universal list of allowed and disallowed phrases. For each synthetic conversation, experts reviewed the history and wrote criteria for the final user message. Asking what kind of help would be useful may earn points in one scenario. Assuming how the person feels or making the decision for them may receive a penalty.
Each criterion has a weight from -10 to +10. At least three experts reviewed every conversation, and a criterion remained only when two agreed and a third did not contradict them. An automated grader, GPT-5.6 Sol, then compares model responses with those expert-written rubrics.
For teams building AI systems, the format carries a practical lesson: safety is not only a blocking layer. It also means gathering context, preserving agency, offering useful next steps and recognizing when the right direction is real-world support. The benchmark separates ten behavioral dimensions, which can reveal differences that one overall score would hide.
Experts and users do not value exactly the same things
OpenAI also ran a separate analysis with 44 adults from 16 countries who had used AI for emotional support or mental-health conversations. Users placed more emphasis on tone and practical next steps. Experts placed more emphasis on collecting context and interpreting ambiguous situations carefully.
That difference is useful. A response can be clinically cautious but still feel cold and unhelpful. It can also sound supportive while moving too quickly toward a conclusion. Looking at both perspectives gives product teams a better chance of avoiding a false choice between safety and a good experience.
A benchmark does not turn a chatbot into a therapist
OpenAI explicitly says ChatGPT is not a substitute for therapy or professional care. MentalHealthBench does not validate a product for clinical use either. Its conversations are synthetic, its criteria reflect expert consensus, and another model performs the grading. Each design choice deserves independent scrutiny.
Even with those limits, publishing the rubrics provides something the field has lacked: a common basis for testing not only what an AI avoids saying, but the quality of the support it tries to provide. Researchers can inspect the method, repeat evaluations and look for cultural, linguistic or clinical gaps.
For organizations putting assistants in front of people during sensitive situations, the message is clear. A content policy is not enough. Teams need to test full conversations, measure behaviors separately and involve specialists in defining what a good response looks like. The value of MentalHealthBench is less about naming a winning model and more about making that discussion concrete and reviewable.
