COMET: Contrastive Motion-Enhanced Temporal Reasoning for Video Multimodal Large Language Models

📅 2026-08-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出COMET框架,通过显式时间表示、外观-运动融合及方向感知优化,解决了视频多模态大语言模型中细粒度运动-时间理解不足的问题。
📝 Abstract
Video multimodal large language models have advanced significantly, yet fine-grained motion-temporal understanding remains fragile. The core bottleneck is not only sparse frame sampling, but also the lack of a complete temporal modeling pipeline for explicitly representing frame-to-frame change, enabling appearance-motion interaction, and optimizing temporal direction sensitivity. We propose COMET, a temporally grounded framework that systematically strengthens video MLLMs through explicit temporal representation, appearance-motion fusion, and direction-aware optimization. Architecturally, COMET introduces a temporal motion branch built on Taylor frame differences and injects its motion evidence into the appearance stream via temporal attention bias-enhanced cross-attention. For optimization, COMET combines temporal prior distillation with a forward-reverse TC-GRPO stage that turns temporal order into a direct learning signal and strengthens the model's use of directional motion patterns encoded by the temporal motion branch. The method achieves consistent overall improvements with a pronounced motion-temporal bias: on Qwen3-VL-8B, action-centric tasks (STAR, SSv2) improve by 4.9% on average, temporal reasoning tasks (NExT-QA, CLEVRER, LLaVA-178K) by 2.1% over BL-GRPO, while static perception tasks (PerceptionTest) remain on par. The same gain pattern also transfers to InternVL2.5-8B, indicating that COMET generalizes across model families.
Problem

Research questions and friction points this paper is trying to address.

Video Multimodal Large Language Models
Temporal Understanding
Frame Sampling
Temporal Modeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Temporal Motion Branch
Taylor Frame Differences
Appearance-Motion Fusion
Direction-Aware Optimization
Temporal Prior Distillation
C
Chenghua Zhu
Guangdong Provincial Key Laboratory of Ultra High Definition Immersive Media Technology, Shenzhen Graduate School, Peking University, Shenzhen, China
Z
Zhaolu Kang
School of Software and Microelectronics, Peking University, Beijing, China
Q
Qifan Shi
School of Future Technology, South China University of Technology, Guangzhou, China
S
Siyan Wu
Aberdeen Institute of Data Science and Artificial Intelligence, South China Normal University, Foshan, China
K
Kehan Jiang
Peking University, Beijing, China
L
Lei Wei
Peking University, Beijing, China
L
Lianyu Hu
Nanyang Technological University, Singapore, Singapore
G
Guangyuan Dong
National University of Singapore, Singapore, Singapore
M
Mingbo Yang
Sun Yat-Sen University, Guangzhou, China
R
Rui Lu
Pingan Technology, Shenzhen, China
Guibo Luo
Guibo Luo
Peking University
medical imagingprivacy computing