Whose Assessment of Distress? Community Perspectives and LLM Alignment on Well-Being Posts

📅 2026-08-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决LLM在评估心理困扰时未能准确反映社区视角的问题,通过社区特定标注和模型对比方法发现现有模型高估轻微困扰情况。
📝 Abstract
Judgments about psychological distress are socially situated: what counts as concerning hinges on community norms around emotional expression, vulnerability, and help-seeking. Yet large language models (LLMs) used for distress detection are typically aligned to a single, undifferentiated standard. How well do these models capture the perspectives of the communities whose language they assess? We address this question through a perspectivist annotation study in which 321 participants provided 9,587 judgments on 1,198 Reddit posts spanning six identity-based communities, yielding community-specific labels. Raters in the contextualized in-group condition show a modest tendency to agree more with their community than uncontextualized out-group raters (OR = 1.18), an effect varying significantly across communities. We then evaluate nine open-weight LLM configurations and four frontier configurations against these labels. Open-weight LLMs systematically over-estimate distress: when communities perceive none-to-mild distress, these models achieve only 31-44% accuracy, predominantly producing false positives. GPT-5 and Gemini 2.5 Pro show the same none-to-mild inflation even when their full-sample over/under rates are mixed, while Claude Opus 4 is more conservative. This pattern does not simply mirror an outsider reading position: uncontextualized out-group human aggregates were nearly symmetric, with 18% over-estimation versus 19% under-estimation. Instead, the models that inflate none-to-mild cases exhibit a distress prior that exceeds both contextualized in-group and uncontextualized out-group human judgments. These findings have implications for equitable AI deployment in mental health contexts, where miscalibrated distress detection may unevenly affect the communities being assessed.
Problem

Research questions and friction points this paper is trying to address.

Distress Detection
Community Norms
Large Language Models
Perspectivist Annotation
Equitable AI Deployment
Innovation

Methods, ideas, or system contributions that make the work stand out.

perspectivist annotation study
community-specific labels
over-estimation of mild distress
equitable AI deployment
💼 Related Jobs
No related jobs found.
A
Andrew Aquilina
School of Computing and Information, University of Pittsburgh
Xiang Lorraine Li
Xiang Lorraine Li
Assistant Professor, University of Pittsburgh
Natural Language ProcessingMachine Learning
Y
Yu-Ru Lin
School of Computing and Information, University of Pittsburgh