CompanionHarm: A Multi-Turn Benchmark for Detecting Harms in Real-World AI Companion Conversations

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决AI伴侣对话中的社会情感伤害检测问题,本文通过构建包含2111个多轮对话的数据集CompanionHarm,并利用多轮对话上下文信息改进现有大语言模型的伤害检测能力。
📝 Abstract
As AI companions become increasingly embedded in everyday life, there is an urgent need to detect harms that emerge in social and emotional human-AI interactions. Yet research in this area is constrained by the lack of real-world, multi-turn conversational datasets for operationalizing and evaluating harms that are relational and contextual. In this work, we introduce CompanionHarm, a publicly available benchmark dataset comprising 2,111 real-world, multi-turn conversations (14,051 utterances) between users and the AI companion Replika. 7,016 AI utterances were annotated independently by three annotators across 13 harmful behavior categories grounded in a taxonomy of AI companion harms, and the dataset includes both aggregated labels and annotator-level labels to support model evaluation and systematic disagreement analysis. Evaluations of seven large language models (LLMs) show that harm detection using multi-turn conversational context outperforms detection based on isolated utterances, although current LLMs still struggle to consistently integrate contextual cues, calibrate harm severity, and interpret relational boundaries. We also find substantial annotator disagreement for context-dependent harmful behaviors, with disagreement varying according to annotators' political affiliation, conversation length, and the utterance's position. Together, CompanionHarm provides a foundation for detecting socio-emotional harms in multi-turn human-AI conversations and for rigorously examining how such harms are interpreted by both humans and LLMs. Our dataset is available at https://github.com/HanMeng2004/CompanionHarm.
Problem

Research questions and friction points this paper is trying to address.

AI companion
harm detection
multi-turn conversations
relational and contextual harms
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-turn conversational dataset
harm detection
contextual cues
annotator disagreement
AI companion
💼 Related Jobs
No related jobs found.
Renwen Zhang
Renwen Zhang
Assistant Professor, Nanyang Technological University
HCIMental HealthSocial SupportHealth CommunicationInterpersonal Communication
Han Meng
Han Meng
National University of Singapore
Human-AI InteractionHuman-Centered NLPComputational Social Science
J
Jian Chai
School of Computing, National University of Singapore
Y
Yuntao Lin
School of Computing, National University of Singapore
Y
Yi-Chieh Lee
School of Computing, National University of Singapore