🤖 AI Summary
This study addresses the challenges of implicit evidence and cross-modal conflicts in multimodal depression assessment by proposing an evidence-centric agent framework. The proposed method transforms implicit feature fusion into explicit evidence deliberation, employing support-challenge branching to arbitrate conflicts while incorporating a risk reflection mechanism to mitigate missed diagnoses. Notably, this framework achieves competitive performance across multiple benchmarks without requiring fine-tuning. By effectively enhancing both interpretability and robustness, this work establishes a novel trustworthy paradigm for multimodal mental health analysis, offering a significant advancement in handling complex multimodal interactions for clinical assessment tasks.
📝 Abstract
Multimodal depression risk assessment requires jointly interpreting textual, acoustic, and visual cues that are often subtle, non-specific, context-dependent, and potentially inconsistent across modalities. Existing multimodal approaches predominantly learn latent representations through feature fusion, leaving the evidence underlying a prediction and the treatment of cross-modal disagreement largely implicit. We propose DepressionAgent, an evidence-centric agentic framework that transforms multimodal depression assessment from implicit feature fusion into explicit evidence deliberation. DepressionAgent first converts textual, acoustic, and visual inputs into modality-specific evidence, and then organizes self-report and behavioral evidence into parallel support--challenge deliberation branches. Cross-modal arbitration explicitly examines agreement and disagreement between the two branches, with conflict reflection revisiting inconsistent assessments before decision making. A subsequent risk reflection mechanism provides an independent textual second opinion for initially low-risk cases to reduce potentially missed risk signals. Without depression-specific supervised training or parameter fine-tuning, DepressionAgent achieves competitive performance on multiple public benchmarks. Extensive ablations, cross-model evaluations, qualitative analyses, and clinician assessments further demonstrate the effectiveness and inspectability of the proposed framework.