CONDUIT: A Unified Residual-Stream Restoration Framework for KV Cache Reuse in Vision-Language Models

📅 2026-09-05
📈 Citations: 0
Influential: 0
📄 PDF
📝 Abstract
Vision-language models (VLMs) often answer new questions about recurring visual content, where reusing the key-value (KV) cache can avoid re-encoding expensive visual prefixes. Exact-prefix reuse, however, fails when the same visual content appears under a changed prefix. Selective recomputation can recover quality under a small visual-token budget, but only when the right stale tokens are refreshed. Raw-attention selection can waste budget on high-attention tokens with small value-norm proxy scores and on query-irrelevant images. To address these failure modes, we propose CONDUIT, a training-free refresh policy that unifies single- and multi-image reuse as residual-stream restoration. Building on norm-weighted attention, CONDUIT ranks cached visual tokens using cached-key query attention and an accessible pre-output cached-value-norm proxy, then applies empirical image-level relevance amplification before one global selection. With one image, the coefficient is one and the rule reduces to intra-image token selection. The method preserves model architecture and weights, adding only a single query-conditioned scoring pass at inference. At a 10% refresh budget, CONDUIT achieves 97.0-99.5% of the corresponding full-prefill five-dataset average across three VLM backbones and leads budgeted methods on average; on the MMLongBench-Doc latency subset, it uses 13.5% of full-prefill FLOPs and achieves a 2.99x time-to-first-token speedup.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
KV Cache Reuse
Prefix Reuse
Selective Recomputation
Attention Selection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unified Residual-Stream Restoration
Norm-Weighted Attention
Empirical Image-Level Relevance Amplification
Training-Free Refresh Policy
🔎 Similar Papers
2024-03-04Computer Vision and Pattern RecognitionCitations: 3
P
Pengan Chen
The Chinese University of Hong Kong
K
Kaisheng Zheng
The Chinese University of Hong Kong
Liang Hong
Liang Hong
PhD at The Chinese University of Hong Kong
L
Lixia Yi
Fudan University
J
Jiyue Jiang
The Chinese University of Hong Kong
J
Jiayang Chen
The Chinese University of Hong Kong
Yixuan Wang
Yixuan Wang
Chinese University of Hong Kong
Machine LearningNeural NetworksBioinformatics
Yimin Fan
Yimin Fan
The Chinese University of Hong Kong
Single-cell genomicsFoundation Models
X
Xinyuan Liu
The Chinese University of Hong Kong
J
Jiayi Li
The Chinese University of Hong Kong
Zhanqiu Zhang
Zhanqiu Zhang
University of Science and Technology of China
Yiwen Guo
Yiwen Guo
Research Scientist
Machine LearningDeep LearningImage Processing
Y
Yu Li
The Chinese University of Hong Kong