Not All Visual Tokens Are Equally Safe to Remove:Consequence-Sensitive Visual Token Compression

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses a critical limitation in conventional visual token compression methods, which disregard the varying real-world costs of different prediction errors in downstream tasks. The authors propose the first consequence-aware compression framework that dynamically allocates token budgets based on task-specific error consequence signals. By offline calibrating error–budget curves and online adjusting token allocation—supporting both token dropping and resolution reallocation—the method accommodates diverse vision-language model architectures. Under identical total computational budgets, it reduces high-risk error rates from 0.300 to 0.133, achieves a 38% reduction in cost-weighted errors under mixed real-world workloads, and lowers inference latency by approximately 21%, substantially outperforming content-driven or uniform allocation strategies.
📝 Abstract
Visual token compression for vision--language models (VLMs) has largely relied on criteria such as attention, redundancy, and uncertainty to maximize average accuracy under a fixed compute budget, implicitly assuming that all errors carry equal cost. However, the consequence of an incorrect prediction on downstream tasks is rarely symmetric: misreading an invoice amount can be far more costly than misclassifying a background color. Motivated by this, we introduce consequence-sensitive visual token compression, which allocates visual computation across requests according to their potential error costs. Our method follows a calibrate-then-allocate procedure, estimating consequence-specific error-budget curves offline and applying the calibrated token budgets online using consequence signals available from question or task information. On a controlled within-task benchmark, high- and low-consequence questions are drawn from the same document images, so content alone cannot reveal which questions are costly to get wrong. In this setting, our method reduces high-stakes errors from 0.300 to 0.133 under the same total token budget, whereas a content-driven allocator performs no better than uniform allocation. Measuring how error rates change with token budget across different cost ratios, we derive an allocation frontier: uniform allocation is optimal when errors are equally costly, and token transfer toward high-consequence questions becomes increasingly beneficial as the cost gap grows. This allocation principle generalizes well across three dense vision-language benchmarks, two budget realization mechanisms (token deletion and resolution reallocation), two VLM architectures, and multiple token selection strategies. On a realistic mixed workload, consequence-sensitive allocation reduces cost-weighted error by 38% while achieving approximately 21% lower latency than full-resolution inference.
Problem

Research questions and friction points this paper is trying to address.

visual token compression
consequence-sensitive
vision-language models
error cost
compute budget
Innovation

Methods, ideas, or system contributions that make the work stand out.

consequence-sensitive compression
visual token allocation
cost-aware inference
vision-language models
error-budget frontier
💼 Related Jobs
No related jobs found.
J
Jingbo Wen
The University of Sydney
L
Liang He
Tongji University
M
Mingyu Cao
University of Surrey
H
Haoyu Wang
Nankai University
M
Minxuan Hu
Cornell University
Kangning Cui
Kangning Cui
Research Assistant Professor of Computer Science, Wake Forest University
Applied MathematicsComputational SustainabilityMedical Imaging
X
Xilu Wang
University of Surrey