MMHBench: A Multi-Perspective Benchmark for Mental Health Understanding in Long-Form Videos

๐Ÿ“… 2026-07-30
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Existing benchmarks for mental health video understanding predominantly rely on coarse-grained classification, which is insufficient to evaluate whether models genuinely possess deep psychological reasoning capabilities. To address this limitation, this work proposes MMHBenchโ€”a multimodal, multi-perspective evaluation benchmark tailored for long-form videos, comprising 268 videos and 2,184 fine-grained questions that span third-person behavioral interpretation and first-person mental state inference. The benchmark introduces an innovative multi-perspective psychological understanding framework, integrating a role-simulation-based Multi-Agent Question Generation (MAQG) mechanism with expert validation to enable precise assessment of modelsโ€™ psychological reasoning abilities. Experiments across 22 state-of-the-art multimodal large language models reveal significant performance gaps, highlighting the continued challenges in this domain.
๐Ÿ“ Abstract
Mental health understanding in long-form videos requires nuanced reasoning over observable behavior, interpersonal context, and latent psychological states. Existing benchmarks largely reduce this task to coarse-grained classification, providing limited insight into whether models truly understand psychological phenomena or rely on superficial correlations. To address this limitation, we introduce MMHBench, a comprehensive multimodal benchmark for multi-perspective mental health understanding, comprising 268 long-form videos and 2,184 carefully curated questions. MMHBench organizes the evaluation into two complementary settings: (1) third-person assessment, consisting of 605 questions that focus on the interpretation of observable behaviors and multimodal evidence, and (2) first-person perspective-taking, comprising 1,579 questions that require perspective-conditioned reasoning to identify the interpretation of the mental state supported by the available multimodal evidence. We propose a Multi-Agent Question Generation (MAQG) framework that simulates diverse social roles to synthesize questions from multiple perspectives. The generated questions are refined through multi-role feedback and iterative optimization, followed by expert-guided verification to ensure quality and validity. Extensive evaluation of 22 representative multimodal large language models (MLLMs), spanning both open-source and leading closed-source models, demonstrates that long-form video mental health understanding remains highly challenging.
Problem

Research questions and friction points this paper is trying to address.

mental health understanding
long-form videos
multimodal benchmark
perspective-taking
psychological reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

MMHBench
Multi-Agent Question Generation
mental health understanding
long-form video
multimodal benchmark
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
Jinpeng Hu
Jinpeng Hu
Hefei University of Technology
natural language processingnamed entity recognitionsummarization
E
Erqiang Wang
Hefei University of Technology
S
Shan Wang
Faculty of Psychology, Beijing Normal University
Zhuo Li
Zhuo Li
The Chinese University of Hong Kong, Shenzhen
Machine LearningNLP
Peipei Song
Peipei Song
University of Science and Technology of China
MultimediaComputer VisionMachine Learning
X
Xun Yang
University of Science and Technology of China
M
Meng Wang
Hefei University of Technology