STAR-Pro: Stage-Wise Token Adaptive Reduction with Progressive Refinement for Efficient Large Vision-Language Models

📅 2026-09-05
📈 Citations: 0
Influential: 0
📄 PDF
📝 Abstract
Large vision-language models (LVLMs) achieve strong multimodal understanding, but the hundreds to thousands of visual tokens they process impose substantial computational overhead, motivating training-free visual token pruning. In this work, we conduct two complementary analyses of visual token pruning. First, we measure the feature-space coverage of tokens retained before cross-modal fusion and find that aggressive pruning discards substantial visual information. Second, we track text-to-visual attention across decoder layers and find that the visual tokens considered important change substantially with depth, making one-shot pruning decisions unreliable. Together, these findings show that effective pruning should preserve broad visual coverage before fusion and progressively refine the retained tokens as cross-modal evidence evolves during fusion. We therefore propose STAR-Pro (STage-Wise Adaptive Token Reduction with Progressive Refinement), a training-free two-stage framework. Its Adaptive Stage applies pivoted QR to construct an over-budget feature-coverage candidate pool, while its Progressive Stage uses evolving text-to-visual attention at selected decoder layers to prune a nested survivor set under a target layer-average token budget. Extensive experiments across seven LVLMs spanning multiple architectures and 18 image and video benchmarks demonstrate the effectiveness of STAR-Pro under aggressive pruning. On LLaVA-Video-7B, STAR-Pro reduces visual tokens by 90.5%, retains 92.7% of baseline performance, and achieves a $2.24\times$ measured inference speedup. Code is available at https://github.com/EasonAI-5589/starpro.
Problem

Research questions and friction points this paper is trying to address.

visual token pruning
large vision-language models
computational overhead
Innovation

Methods, ideas, or system contributions that make the work stand out.

Stage-Wise Adaptive Token Reduction
Progressive Refinement
Visual Token Pruning
Cross-Modal Fusion
Text-to-Visual Attention
Yichen Guo
Yichen Guo
Master student in Nanyang Technological University
T
Tinghao Wang
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University
Qizhe Zhang
Qizhe Zhang
School of Computer Science, Peking University
Vision Language ModelComputer VisionMachine Learning
L
Lingbei Meng
De Artificial Intelligence Lab
Y
Yuan Zhang
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University
Jiajun Cao
Jiajun Cao
Ph.D. Student, Peking University
MLLMComputer Vision
H
Hao Jiang
University of Electronic Science and Technology of China
C
Chenwei Wu
University of Michigan, Ann Arbor
J
Jixian Wu
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University
Sixiang Chen
Sixiang Chen
The Hong Kong University of Science and Technology (Guangzhou)
Computer VisionImage RestorationAIGCMLLM
T
Tao Luo
University of Electronic Science and Technology of China
H
Hongyang Cheng
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University
K
Kai Tang
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University
C
Chenxi Li
De Artificial Intelligence Lab
R
Renyuan Li
University of Electronic Science and Technology of China
Xiande Huang
Xiande Huang
DAIL Tech
Artificial IntelligenceMachine LearningTrustworthy AIMedical AgentsReinforcement Learning
Wenya Wang
Wenya Wang
Nanyang Technological University
Deep LearningKnowledge ReasoningNatural Language ProcessingSentiment Analysis
Shanghang Zhang
Shanghang Zhang
Peking University
Embodied AIFoundation Models