Video-HolmesV2: Can MLLMs Reason with Spatio-Temporal Audio-Visual Evidence in Long Videos?

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了长视频中多模态大语言模型跨模态推理能力不足的问题,通过提出Video-HolmesV2基准和音频-文本引导的令牌压缩框架来增强模型对时空视听证据的理解与利用。
📝 Abstract
Multimodal Large Language Models have demonstrated impressive video understanding, yet their ability to reason over long-form narratives is often masked by visual-centric evaluations and inefficient context processing. Existing benchmarks over-rely on visual heuristics while marginalizing auditory cues, effectively reducing models to "silent observers" that bypass genuine cross-modal reasoning. Moreover, standard dense sampling creates an evidence-context trade-off: increasing frames to capture evidence inevitably leads to attention distraction and token explosion. To bridge these gaps, we present Video-HolmesV2, a novel benchmark designed for Deep Audio-Visual Coupling. Unlike previous works, it enforces an Evidence-Based Evaluation, requiring models to justify answers with precise spatio-temporal audio-visual evidence, thereby reducing confounding effects of guessing and hallucinated evidence. To support this, we introduce: (1) a Multi-Model Cross-Verification pipeline to ensure task rigor; (2) a Spatio-temporal Evidence-Aware Metric for fine-grained calibration. Furthermore, we propose an Audio-Text Guided Token Compression framework. By fusing task intent with auditory anchors, our method distills high-value reasoning cues to mitigate long-context noise. In our evaluation, even strong proprietary models achieve below 60% accuracy, while our approach outperforms comparable open-source omni-models.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
Spatio-Temporal Audio-Visual Evidence
Long Videos
Cross-Modal Reasoning
Evidence-Context Trade-off
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evidence-Based Evaluation
Multi-Model Cross-Verification
Spatio-temporal Evidence-Aware Metric
Audio-Text Guided Token Compression
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Zhaoyang Wei
Zhaoyang Wei
University of Chinese Academy of Sciences
Computer visionPoint PromptPointly SupervisionWeakly SupervisionInteractive Perception
Z
Zipeng Wang
University of Chinese Academy of Sciences, China
Y
Yushe Cao
Tsinghua University, China
C
Chenhui Qiang
Tencent, China
S
Shuaibing Cheng
University of Chinese Academy of Sciences, China
Xuesong Yang
Xuesong Yang
NVIDIA
Machine LearningDeep LearningNatural Language ProcessingSpeech Signal Processing
S
Sen Nie
University of Chinese Academy of Sciences, China
Bowen Jiang
Bowen Jiang
University of Pennsylvania, Microsoft Corporation
Artificial IntelligencePost-trainingPersonalizationMultimodality
Wenchao Ding
Wenchao Ding
Tenure-track Associate Professor, Fudan University
RoboticsMotion PlanningAutonomous NavigationDecision Making
Y
Yanchao Hao
Tencent, China
Z
Zheng Wei
Tencent, China
X
Xuehui Yu
Tencent, China
Z
Zhenjun Han
University of Chinese Academy of Sciences, China