HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决通用健康基准难以评估特定领域性能的问题,通过筛选和验证建立了专门针对心理健康领域的HealthBench-Psych基准。
📝 Abstract
General-purpose health benchmarks increasingly anchor claims about LLM medical performance, but they are not always resolved by clinical specialty, making domain-specific performance hard to isolate. Mental health is of acute public-health concern as millions of people turn to LLMs for psychological support, and most existing evaluations are bespoke academic benchmarks that are difficult to integrate into developer workflows. We introduce HealthBench-Psych and HealthBench-Psych-Hard. We screened HealthBench's 5,000 physician-rubric conversations for mental-health relevance with a transparent LLM-applied rubric, then validated the subset through two rounds of blinded clinician review with concealed known-exclude controls, yielding 610 conversations (12.2% of the corpus). Evaluating 20 frontier and open models under a cross-vendor panel of three LLM judges, we find a statistically tied frontier cluster, measurable refusal behavior in two models, and near-identical rankings across judges ($τ\ge 0.92$). We release the subset, pipeline, model responses, grades, and analysis code as a reusable resource.
Problem

Research questions and friction points this paper is trying to address.

mental health
large language models
health benchmarks
clinical specialty
Innovation

Methods, ideas, or system contributions that make the work stand out.

mental health benchmark
LLM-applied rubric
clinician review
model performance evaluation
reusable resource
🔎 Similar Papers
No similar papers found.
M
Matthew Flathers
Division of Digital Psychiatry, Beth Israel Deaconess Medical Center, Boston, MA
P
Phuong Anh Nguyen
Division of Digital Psychiatry, Beth Israel Deaconess Medical Center, Boston, MA
J
Jill Noorily
Division of Digital Psychiatry, Beth Israel Deaconess Medical Center, Boston, MA
J
Julian Herpertz
Division of Digital Psychiatry, Beth Israel Deaconess Medical Center, Boston, MA
M
Meiting Chen
Division of Digital Psychiatry, Beth Israel Deaconess Medical Center, Boston, MA
J
Jasreen Multani
Division of Digital Psychiatry, Beth Israel Deaconess Medical Center, Boston, MA
Samuel Powell
Samuel Powell
Department of Psychiatry, Beth Israel Deaconess Medical Center, Boston, MA
M
Mason Granof
Division of Digital Psychiatry, Beth Israel Deaconess Medical Center, Boston, MA
M
Mark Kalinch
Division of Digital Psychiatry, Beth Israel Deaconess Medical Center, Boston, MA
John Torous
John Torous
Harvard Medical School / Beth Israel Deaconess Medical Center
Clinical InformaticsDigital PhenotypingSmartphonesMental HealthPsychiatry