K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为评估大型语言模型在高风险心理健康对话中的安全性,开发了K-Bench基准测试,通过临床校准和多情景模拟来评价模型表现。
📝 Abstract
% !TEX root = ../main.tex People increasingly use large language models (LLMs) for mental health support, yet their safety in evolving, high-risk conversations remains poorly characterised. We developed K-Bench, a clinician-calibrated, protected benchmark evaluating 125 model configurations representing 33 base models from 14 providers across a fixed cohort of 200 multi-turn vignettes involving suicide, self-harm, domestic violence, substance misuse, and no-risk presentations. Synthetic patient conversations showed substantial distributional overlap with real human-AI conversations. A frozen GPT-4o judge achieved 94.2% exact agreement with clinician consensus across 6,751 eligible item comparisons from 151 clinician-rated transcripts. Leading models combined strong supportive conversation with combined-risk scores above 95, whereas risk exploration exposed substantial variation among lower-performing configurations. Therapeutic prompting produced configuration-specific gains concentrated among weaker models, while elevated reasoning produced no average improvement. K-Bench combines broader clinical coverage and configuration-scale comparison with a continuously updated public leaderboard whose operational test materials are protected from direct optimisation. The leaderboard is available at www.k-bench.ai.
Problem

Research questions and friction points this paper is trying to address.

large language models
mental health
high-risk conversations
safety
Innovation

Methods, ideas, or system contributions that make the work stand out.

clinician-calibrated benchmark
large language models
high-risk mental health conversations
GPT-4o judge
public leaderboard
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Laura M. Vowels
School of Psychology, University of Roehampton, London, United Kingdom
M
Matthew J. Vowels
Kivira Health, United Kingdom
S
Shivali Sharma
School of Psychology, University of Roehampton, London, United Kingdom
A
Apoorv Jha
Kivira Health, United Kingdom
R
Rehnuma Choudhury
School of Psychology, University of Roehampton, London, United Kingdom
W
Wasseem El Sarraj
University of Hertfordshire, Hatfield, United Kingdom
R
Rachel Francois-Walcott
University of Surrey, Guildford, United Kingdom
A
Aruba Hussain
School of Sport, Psychology and Social Sciences, University of Bedfordshire, Luton, United Kingdom
S
Sarah Ingram
Tavistock Relationships, London, United Kingdom
A
Angela Loulopoulou
School of Psychology, University of Roehampton, London, United Kingdom
A
Adva Segal
InsideOut, United Kingdom
E
Elena Volkova
School of Psychology, University of Roehampton, London, United Kingdom