OccamToken: Efficient VLM Inference with Training-Free and Budget-Adaptive Token Pruning

📅 2026-05-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the high computational and memory costs incurred by long visual token sequences in vision-language model (VLM) reasoning, where existing fixed-budget token pruning methods struggle to balance input redundancy and query relevance. The authors propose OccamToken, a training-free dynamic token pruning framework that abandons conventional absolute importance scoring. Instead, it introduces register tokens as stable references and leverages a register attention mechanism coupled with relative evidence testing to achieve image-adaptive redundancy pruning and query-adaptive relevance pruning. Supporting dynamic budgets, OccamToken retains only about 40 tokens (1.4% of the original sequence) on mainstream VLMs such as LLaVA-NeXT while preserving over 93% of the original accuracy, substantially improving the trade-off between efficiency and performance.
📝 Abstract
Vision-language models (VLMs) rely on long visual token sequences for visual understanding, making the prefill stage expensive in both computation and memory. Most existing pruning methods follow an absolute-ranking paradigm, assigning importance scores to visual tokens and retaining a fixed top-K subset. In this work, we argue that this paradigm is fundamentally brittle: attention sinks distort token importance rankings, while image redundancy and query-dependent visual evidence make fixed token budgets unreliable across inputs. We propose OccamToken, a training-free framework that replaces absolute token ranking with register-anchored relative evidence testing. Instead of asking which tokens are globally important, OccamToken evaluates whether a visual token provides information beyond a register-based reference. Our key insight is that register tokens naturally absorb low-information attention patterns, making them a stable reference for identifying genuinely informative visual evidence. Based on this principle, OccamToken performs both image-adaptive redundancy pruning and query-adaptive relevance pruning through dynamic thresholds derived from register attention. Across LLaVA-NeXT, LLaVA-v1.5, and Qwen3-VL, OccamToken consistently improves the accuracy-efficiency trade-off without additional training. Notably, on LLaVA-NeXT, it reduces 2,880 visual tokens to approximately 40 while preserving over 93% of the original accuracy, enabling stable visual token compression even in the extreme 1.4% retention regime.
Problem

Research questions and friction points this paper is trying to address.

vision-language models
token pruning
inference efficiency
visual redundancy
query-dependent relevance
Innovation

Methods, ideas, or system contributions that make the work stand out.

token pruning
vision-language models
training-free
adaptive thresholding
register-anchored
💼 Related Jobs
No related jobs found.
Geng Li
Geng Li
Peking University
Guohao Chen
Guohao Chen
South China University of Technology
Transfer LearningDomain AdaptationTest-Time Adaptation
T
Ting Chen
Nanyang Technological University (NTU)
S
Shilin Shan
Nanyang Technological University (NTU)
K
Kuangji Zuo
Nanyang Technological University (NTU)
B
Bofan Lyu
Nanyang Technological University (NTU)
T
Tuo An
Nanyang Technological University (NTU)
Gen Li
Gen Li
Postdoctoral Research Fellow, Nanyang Technological University
Embodied AIComputer VisionRoboticsArtificial Intelligence
Jianfei Yang
Jianfei Yang
Assistant Professor, Director of MARS Lab, Nanyang Technological University
Physical AIEmbodied AIMultimodal AI