VideoMM: Adaptive Macro-Micro Inference for Efficient Video MLLMs

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决长视频理解中视觉令牌过多的问题,VideoMM通过自适应宏微观推理方法,有效筛选语义相关区域并仅在必要时进行详细理解,提高效率和准确性。
📝 Abstract
Scaling Multimodal Large Language Models (MLLMs) to long-form video understanding is bottlenecked by the explosion of visual tokens, which saturates context windows and incurs prohibitive costs. Current solutions predominantly rely on auxiliary models for token reduction but face a fundamental dilemma: lightweight encoder-driven approaches often overlook critical semantic information, whereas heavyweight MLLM-driven reduction negates the efficiency gains. {In this work, we identify a more fundamental inefficiency underlying this dilemma: while fine-grained visual details are essential for detailed understanding, they are largely redundant for the preliminary task of selecting semantically relevant regions. } Motivated by this, we introduce \textbf{VideoMM}, which marks a paradigm shift from model-centric downsizing to adaptive perceptual granularity. Specifically, our framework {decouples selection from reasoning} by executing semantic filtering on a cost-effective \textit{Macro Proxy} (derived from downscaled frames), and projecting the selected regions onto high-fidelity \textit{Micro Tokens} for detailed understanding only when necessary. Extensive evaluations show that VideoMM significantly outperforms existing solutions. It achieves a 6.13$\times$ speedup and a 7.4\% accuracy gain over full-context baselines on LongVideoBench, and further accelerates inference by 2.73$\times$ over current leading methods, establishing a highly scalable paradigm for long-video understanding. Our code is available at: https://github.com/adfh917k/VideoMM.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
long-form video understanding
visual tokens
context windows
auxiliary models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive Macro-Micro Inference
Semantic Filtering
Macro Proxy
Micro Tokens
🔎 Similar Papers
2024-02-20International Conference on Machine LearningCitations: 30
Haoyu Guo
Haoyu Guo
Shanghai AI Lab
Computer Vision3D Vision
Y
Yuan Feng
School of Computer Science, USTC; Data Darkness Lab, MIRACLE Center, Suzhou Institute for Advanced Research
Junlin Lv
Junlin Lv
USTC
machine learning system
Mingjun Xiao
Mingjun Xiao
University of Science and Technology of China
Mobile ComputingCrowdsensingMobile Social NetworkVechular Network
S
S Kevin Zhou
School of Biomedical Engineering, University of Science and Technology of China; Data Darkness Lab, MIRACLE Center, Suzhou Institute for Advanced Research
X
Xike Xie
School of Biomedical Engineering, University of Science and Technology of China; Data Darkness Lab, MIRACLE Center, Suzhou Institute for Advanced Research