OpenAI releases MentalHealthBench, an open benchmark co-developed with 80+ licensed clinicians to evaluate AI in realistic mental health conversations
What happened
OpenAI 放出了一个叫 MentalHealthBench 的开源评测集,专门看 AI 在心理健康对话里表现怎么样。这个评测集是和 22 个国家、80 多位持证心理医生和精神科医生一起做的,覆盖了从日常情绪倾诉到紧急危机干预的各种场景,还分了成人、青少年、照顾者等不同角色。它不只看模型有没有触发安全红线,更会检查模型会不会主动追问背景、有没有尊重用...
Coverage
Follow the reports to see the story from different sides.
- OpenAI NewsPickOpenAI releases MentalHealthBench, an open benchmark co-developed with 80+ licensed clinicians to evaluate AI in realistic mental health conversations
OpenAI open-sourced MentalHealthBench, a benchmark built with over 80 licensed psychologists and psychiatrists across 22 countries. It tests AI on realistic mental health conversations ranging from everyday stress to emergencies, covering adults, teens, and caregivers. The eval goes beyond safety filters: it checks whether models seek context, preserve user agency, and offer actionable guidance when appropriate. OpenAI stresses ChatGPT isn't a substitute for therapy, but the benchmark tracks progress on empathy and steering people toward real-world support. The paper and benchmark are publicly available.