From Saliency to Discriminability: Rank-Preserving Visual Token Pruning for VLM Rerankers

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对视觉-语言模型在重排序时的效率问题,提出了一种基于归一化注意力熵的选择性剪枝方法RaDiCal,有效减少了计算量并保持了性能。
📝 Abstract
Large vision-language models used as listwise rerankers must jointly process visual tokens from tens of candidates per query, making token pruning essential for practical deployment. Existing pruning methods retain tokens by attention saliency, yet we show that saliency is systematically misaligned with ranking contribution: visually prominent tokens often capture order-neutral patterns shared across candidates. This mismatch is layer-dependent: saliency becomes informative only where attention is concentrated, and normalized attention entropy diagnoses the reliability shift (Pearson r=0.87). We propose RaDiCal (Rank-Discriminative Calibration), a training-free framework that uses normalized attention entropy to decide when saliency can be trusted, fusing it with an attention-free rank-discriminative prior and selecting pruning layers from the same trust landscape. Across three retrieval benchmarks and multiple VLM architectures, RaDiCal matches Dense MRR@10 on Flickr30K and surpasses it on MSCOCO at a 20% token budget, ranks first among all pruning methods on FashionIQ, and holds within 1.2 pp on Flickr30K and MSCOCO at 10% retention. It cuts FLOPs by 39--45% and delivers 1.28--1.45$\times$ measured speedups across two VLM architectures without dataset-specific retuning.
Problem

Research questions and friction points this paper is trying to address.

visual tokens
saliency
ranking contribution
attention saliency
token pruning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Normalized Attention Entropy
Rank-Discriminative Prior
Token Pruning
Siyi Liu
Siyi Liu
Hong Kong University of Science and Technology (Guangzhou)
Recommender SystemsInformation Retrieval
H
Hanjun Yang
The Hong Kong University of Science and Technology (Guangzhou)
C
Chenchen Zhang
The Hong Kong University of Science and Technology (Guangzhou)
X
Xiaorong Zhu
The Hong Kong University of Science and Technology (Guangzhou)
X
Xinyu Zuo
Tencent Yuanbao
L
Lisheng Duan
Tencent Yuanbao
H
Haijin Liang
Tencent Yuanbao
J
Jin Ma
Tencent Yuanbao
J
Junfu Pu
ARC Lab, Tencent
Yongqi Zhang
Yongqi Zhang
Assistant Professor in HKUST(GZ)
Graph learningDrug discoveryDeep learning