StepPrune: Adaptive Sequential Visual Token Selection across Multimodal Large Language Models

📅 2026-09-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
针对多模态大语言模型中视觉token计算量大的问题,提出StepPrune方法,通过自适应顺序决策过程选择视觉token,减少推理时间。
📝 Abstract
Visual prefixes account for a major portion of the per-layer computation in multimodal large language models (MLLMs), making visual-token pruning a direct approach to accelerating inference. Existing top-K methods typically evaluate tokens independently and apply a uniform budget to all inputs, overlooking both selection-dependent interactions and variations in visual complexity across samples. In contrast, we propose StepPrune, which formulates visual-token pruning as an adaptive sequential decision process. Conditioned on previously selected tokens and textual context, StepPrune progressively constructs the retained subset and automatically determines its size through a learned STOP action. During training, a variance-preserving noise gate provides a differentiable surrogate for the discrete selection process, whereas during inference, unselected tokens are physically removed before language-model prefill. A grouped selection mechanism further extends StepPrune to high-resolution inputs. Experiments across LLaVA-1.5, LLaVA-NeXT, Qwen2.5-VL, and InternVL3 show that StepPrune achieves the best average normalized performance retention across all evaluated pruning rates on LLaVA-1.5, Qwen2.5-VL, and InternVL3, while remaining competitive on the substantially longer AnyRes prefixes of LLaVA-NeXT. On LLaVA-1.5, StepPrune retains 94.6% of the full-prefix normalized performance while pruning 88.9% of the visual tokens. At a mean retained count of 64, StepPrune reduces prefill latency from 59.95 ms to 40.05 ms, corresponding to a 1.50x prefill speed-up.
Problem

Research questions and friction points this paper is trying to address.

Visual Token Pruning
Multimodal Large Language Models
Adaptive Sequential Decision Process
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive Sequential Decision Process
Variance-Preserving Noise Gate
Grouped Selection Mechanism
💼 Related Jobs
No related jobs found.
H
Hansen Zhang
Shenzhen University of Advanced Technology, Shenzhen, China
L
Landi He
Shenzhen University of Advanced Technology, Shenzhen, China; School of Computer Science, Nanjing University, Nanjing, China
Mingde Yao
Mingde Yao
The Chinese University of Hong Kong <-- USTC
computational photographygenerative models
L
Lijian Xu
Shenzhen University of Advanced Technology, Shenzhen, China