MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models

📅 2026-06-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing vision-language models struggle to capture fine-grained motion details in video understanding, as they primarily focus on static semantics and high-level event logic. To address this limitation, this work proposes a scalable approach that, for the first time, effectively transfers motion priors from video diffusion models to vision-language models without requiring additional parameters, architectural modifications, or external tools. The method leverages two parameter-free modules—Motion-sensitive Head Selection (MHS) and Motion-salient Text Token Identification (MTTI)—which enhance motion awareness through an attention alignment mechanism. Experimental results demonstrate that the proposed approach significantly outperforms state-of-the-art models on two motion-centric video understanding benchmarks, achieving consistent improvements particularly on motion-related evaluation metrics.
📝 Abstract
The new era has witnessed a remarkable capability to extend Vision-Language Models (VLMs) for tackling tasks of video understanding. While current VLMs excel at event- or story-level understanding, their ability to capture fine-grained motion details remains limited, primarily due to their focus on high-level static semantic structures and macro-event logic. In contrast, Video Diffusion Models (VDMs) are adept at modeling dynamic motion patterns, benefiting from large-scale video data and the intrinsic requirement of temporal generation. In this paper, we introduce MotionEnhancer, a novel approach that leverages motion priors distilled from a powerful video diffusion model as auxiliary supervision to enhance the motion understanding capability of a VLM via attention alignment. MotionEnhancer comprises two simple parameter-free modules, Motion-sensitive Head Selection (MHS) and Motion-salient Text Token Identification (MTTI), to directly extract and optimize motion-related attentions from the VDM in a computation-only manner. MotionEnhancer provides a scalable solution for motion understanding without additional training parameters, modifications to existing architectures, or tool calling. Extensive experiments demonstrate that MotionEnhancer can achieve consistent improvements over state-of-the-art VLMs on two motion-level video understanding benchmarks, especially on motion-related metrics.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
video understanding
fine-grained motion
motion understanding
Video Diffusion Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Video Diffusion Models
Vision-Language Models
Motion Understanding
Attention Alignment
Parameter-free Enhancement
🔎 Similar Papers
No similar papers found.
Y
Yifan Xu
School of Computer Science and Engineering, Beihang University
C
Chao Zhang
Beijing Digital Native Digital City Research Center
R
Ruifei Ma
Beijing Digital Native Digital City Research Center
F
Fei Gao
Beijing Digital Native Digital City Research Center
Zhifei Yang
Zhifei Yang
Peking University
3D GenerationGenerative Models
Jiaxing Qi
Jiaxing Qi
BUAA
AIOpsSoftware EngineeringData MiningAI4Science
Zhipeng Chen
Zhipeng Chen
Phd research scholar, Beijing Jiaotong University
information securitydigital forensicssignal processing