Multi-Image Visual Token Pruning in Large Visual Language Models

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决大型视觉语言模型处理多图像序列时的计算和上下文长度限制问题,提出了一种无需训练的自适应视觉令牌剪枝框架AVTP,并通过实验证明其有效性和鲁棒性。
📝 Abstract
With the growing demand for processing multiple image sequences in real-world applications, various visual token pruning methods have emerged to mitigate the computational and context length constraints faced by Large Vision Language Models (LVLMs). However, most existing pruning approaches rely on static strategies that struggle to adapt across different architectural LVLMs and multi-image scenarios, and are additionally constrained by their dependence on attention computations that are incompatible with efficient techniques like FlashAttention. To address these limitations, we propose a training-free, Adaptive Visual Token Pruning (AVTP) framework, applicable to diverse LVLM architectures. We strategically determine pruning layers based on empirical analysis of visual attention distributions across various LVLMs, and implement adaptive pruning ratios in multi-image contexts where images of higher importance retain proportionally more tokens. We conduct extensive experiments across different LVLMs to demonstrate the effectiveness and robustness of AVTP. Specifically, Qwen3VL-8B achieves 2 times inference speedup while maintaining 96.1\% of its original accuracy on multiple multi-image benchmarks, InternVL3.5-8B retains 94.1\% accuracy, and LLaVA-OV-7B even exceeds its original baseline performance. Our code is available at \href{https://github.com/zry13/AVTP}{this link}.
Problem

Research questions and friction points this paper is trying to address.

Multi-Image Sequences
Large Vision Language Models
Visual Token Pruning
Computational Constraints
Context Length
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive Visual Token Pruning
Training-free Framework
Multi-Image Contexts
Dynamic Pruning Ratios
Efficiency and Accuracy
💼 Related Jobs
No related jobs found.
R
Rongyang Zhang
State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China
Chengqiang Lu
Chengqiang Lu
USTC
C
Cong Li
Xiaohongshu Inc.
H
Hongchao Gu
State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China
T
Tingjia Shen
State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China
X
Xuyang Zhi
State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China
Q
Qimeng Wang
Xiaohongshu Inc.
Yan Gao
Yan Gao
XiaoHongShu Inc
机器学习
Y
Yi Wu
Xiaohongshu Inc.
Yao Hu
Yao Hu
浙江大学
Machine Learning
H
Hao Wang
State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China
Enhong Chen
Enhong Chen
University of Science and Technology of China
data miningrecommender systemmachine learning