CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional Companionship

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses three key limitations in current evaluations of AI emotional companionship systems: reliance on handcrafted simulated scenarios, reduction of empathy to a single scalar score, and neglect of evaluator bias. To overcome these issues, the authors propose the first bilingual, interactive benchmark grounded in real user data and psychological theory. The framework employs a disclosure-gating mechanism to dynamically control dialogue states and evaluates agents across ten core competencies—four of which, such as tolerance for ambiguity and calibrated challenging, are explicitly assessed for the first time—derived from 25 theoretical constructs. Integrating a real-data-driven user simulator, item response theory, and cross-family review, the method mitigates evaluation bias. Experiments on 28 AI agents reveal that aggregate scores obscure critical capability disparities, with emotion regulation and calibrated challenging consistently weak, and role-playing agents performing worst. The release includes 500 parallel Chinese–English dialogues and code, yielding highly reproducible rankings (ρ = 0.996/0.953).
📝 Abstract
LLM companions are deployed at scale in personally consequential settings, yet poorly evaluated. Existing benchmarks use hand-authored scenarios and prompted simulators, aggregate empathy into one score, and overlook judge biases such as same-family favoritism and scale drift. We introduce CompanionBench, an interactive bilingual benchmark. To our knowledge, it is the first companion benchmark to ground both its scenarios and a trained user simulator in de-identified real-world data. A hidden disclosure gate branches each persona's trajectory on the agent's own behavior, controlling the interaction state space without scripting dialogue. We operationalize ten capabilities derived from 25 theories across psychology and counseling, four of them not graded explicitly by prior work: holding ambiguity, selfobject responsiveness, positive resonance and calibrated challenge. Agents are assessed on two complementary axes: a subjective ten-capability rubric and a deterministic measure of whether deeper disclosure was earned. A cross-family panel dilutes same-family favoritism; an Item Response Theory model separates agent quality from judge severity. Theory fixes what to measure and how personas are structured; real data supply events, history, and profiles -- coverage from theory, authenticity from data. Rankings are reproducible in both languages (rho = 0.996 ZH / 0.953 EN). Evaluating 28 agents reveals capability-level differences obscured by aggregate scores. Emotion regulation and calibrated challenge remain common weaknesses; holding ambiguity discriminates most. Role-play agents rank near the bottom: immersion does not imply relational competence. Across agents, the dominant failure mode is substituting surface warmth for substantive relational support. We will release 500 Chinese-English parallel pairs and the evaluation code.
Problem

Research questions and friction points this paper is trying to address.

AI Emotional Companionship
Benchmark Evaluation
Judge Bias
Empathy Assessment
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

CompanionBench
theory-anchored evaluation
real-world grounded simulator
hidden disclosure gate
relational competence
🔎 Similar Papers
No similar papers found.