Aggregating Visual Information with Optimal Transport for VideoLM Token Compression

📅 2026-08-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为减少视频语言模型中视觉token的负担,本文提出使用最优传输方法(AVIOT)压缩视频序列,同时保持视觉信息并适应任务和空间需求。
📝 Abstract
Video language models process videos as dense visual-token sequences with substantial representational redundancy. Compressing these sequences is therefore essential for reducing the visual-token burden on language-model decoding. The central challenge is to preserve visual information dispersed across frames under such compression. To this end, we introduce Aggregating Visual Information with Optimal Transport (AVIOT), which casts video token compression as transporting a dense empirical measure of frame observations onto a compact target measure. The resulting source-to-target coupling induces a distribution over source observations for each target support, directly specifying how the compressed video representation is constructed. We further adapt this construction along task and spatial axes. Question conditioning modulates the transport cost between source frames and target supports, while influencing how many supports are allocated to each temporal segment, thereby directing representation capacity toward question-relevant content. At multiple spatial granularities, AVIOT computes region-specific temporal transport plans and adaptively fuses the representations they yield, allowing different regions within the same compact representation to draw from different moments. Evaluations across varying compression ratios show that AVIOT matches or outperforms the uncompressed baseline on multiple video-understanding benchmarks while retaining strong performance at higher compression ratios.
Problem

Research questions and friction points this paper is trying to address.

Visual Information
Token Compression
Video Language Models
Representation Redundancy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Optimal Transport
Token Compression
Video Language Models
Representation Redundancy
Question Conditioning
W
Wenti Yin
Key Laboratory of Image Processing and Intelligent Control, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology
Xiaotian Han
Xiaotian Han
Research Scientist, OpenAI
Machine learningComputer VisionMultimodalGenAILLM
Junyuan Shang
Junyuan Shang
Baidu NLP
Deep LearningNatural Language ProcessingHealthcare
Y
Yuchen Ding
Baidu, Inc
Shuohuan Wang
Shuohuan Wang
Baidu
Natural Language ProcessingDeep Learning
Dianhai Yu
Dianhai Yu
Baidu
Deep LearningNatural Language ProcessingMachine LearningArtificial intelligence
C
Changxin Gao
Key Laboratory of Image Processing and Intelligent Control, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology
Nong Sang
Nong Sang
Huazhong University of Science and Technology
Computer Vision and Pattern Recognition