BALMS: Benchmarking Agentic LLMs for Longitudinal Mental Health Sensing

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决心理健康评估的连续监测问题,本文提出BALMS基准,利用大语言模型处理可穿戴设备数据,进行长期心理状态预测与解释。
📝 Abstract
Mental health assessment relies on episodic self-report scales, which convert subjective states such as stress into numerical scores but provide only sparse snapshots of wellbeing. Wearable devices offer longitudinal behavioral and physiological signals for continuous, low-burden monitoring. Recent LLM-driven personal-health agents enable natural language queries over wearable signals, but mainly handle short-term, retrieval-based lookups (e.g., highest step count over a week). They do not evaluate whether agents can reason over long-term signals to predict wellbeing scores paired with evidence-grounded rationales. To address this gap, we introduce BALMS, the first systematic benchmark of LLM-based agentic systems for longitudinal mental health sensing. BALMS spans 3 real-world longitudinal datasets, 2 task families (closed-form wellbeing-score prediction and rationale generation auto-graded by an LLM-as-Judge), 3 agentic paradigms evaluated across 5 open- and closed-source LLM backbones. We find that zero-shot agents rarely outperform a simple mean baseline, except with stronger backbones or compact, semantically meaningful features. Chain-of-thought prompting improves reasoning-oriented backbones, but does not guarantee temporal grounding or numerical correctness. Together with more analysis on efficiency and temporal scaling, BALMS highlights the need for longitudinal mental health agents that selectively retrieve history, ground temporal evidence, and reason over interpretable behavioral features.
Problem

Research questions and friction points this paper is trying to address.

mental health assessment
longitudinal signals
wellbeing prediction
evidence-grounded rationales
agentic LLMs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Longitudinal Mental Health Sensing
Agentic LLMs
Chain-of-thought Prompting
Temporal Evidence Grounding
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yu Yvonne Wu
Dartmouth College
A
Arvind Pillai
Dartmouth College
Yuliang Chen
Yuliang Chen
University of California, San Diego
Self-Supervised LearningMultimodal Learning
Y
Yuwei Zhang
University of Cambridge
Sudarshan Regmi
Sudarshan Regmi
Dartmouth College
Machine Learning
T
Tess Z. Griffin
Dartmouth College
M
Michael V. Heinz
Dartmouth College
L
Lisa A. Marsch
Dartmouth College
Nicholas C. Jacobson
Nicholas C. Jacobson
Dartmouth College
Digital PhenotypingDigital InterventionsArtificial IntelligenceMental HealthChatbots
A
Andrew Campbell
Dartmouth College