FOVEA: Focused On-Demand Visual Evidence Adaptation for Cache-Friendly Multimodal Speculative Decoding

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出FOVEA方法,通过动态调整视觉证据以适应不同任务和解码阶段的需求,从而提高多模态投机解码效率。
📝 Abstract
Multimodal speculative decoding accelerates vision-language models by allowing a lightweight draft model to propose candidate tokens for parallel verification by a larger target model. Existing methods typically condition the drafter on a fixed visual interface, such as a predefined visual-token budget or a static compressed representation. However, our controlled visual-budget analysis shows that visual demand varies substantially across tasks and decoding stages, which means more visual input is not always beneficial. Actually, insufficient evidence may weaken visual grounding, while excessive context adds overhead and may disrupt drafting. We propose FOVEA (Focused On-demand Visual Evidence Adaptation), a cache-friendly approach that builds a reusable visual memory and dynamically retrieves a bounded subset for a draft state. A cumulative-mass rule determines both how many and which entries are selected. The selected entries are aggregated into a visual readout and fused with the current draft hidden state through a lightweight gated residual correction. Rather than inserting visual tokens into the autoregressive context, the correction modifies only the representation passed to the language-model head. Experiments across multiple vision-language backbones and multimodal benchmarks show that FOVEA improves draft acceptance and end-to-end decoding speed, achieving up to $2.13\times$ speedup over autoregressive decoding. These results demonstrate that state-conditioned evidence retrieval is an effective alternative to reusing a fixed visual representation throughout multimodal generation.
Problem

Research questions and friction points this paper is trying to address.

multimodal speculative decoding
visual demand
fixed visual interface
Innovation

Methods, ideas, or system contributions that make the work stand out.

FOVEA
Dynamic Visual Evidence Retrieval
Cumulative-mass Rule
Gated Residual Correction
Cache-friendly
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
H
Hengjie Zhu
Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences
D
Dayan Wu
Institute of Information Engineering, Chinese Academy of Sciences
Zihao Zhang
Zihao Zhang
天津大学
计算机视觉
X
Xinze Liu
Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences
J
Jingxuan Yu
School of Cyber Science and Engineering, Southeast University
Peng Fu
Peng Fu
Institute of Information Engineering, Chinese Academy of Sciences
Natural Language Processing
Zheng Lin
Zheng Lin
Institute of Information Engineering, CAS
NLP
Weiping Wang
Weiping Wang
School of Information Science and Engineering, Central South University
Computer NetworkNetwork Security
Ding Wang
Ding Wang
Institute of Automation, Chinese Academy of Sciences
BioinformaticsMLSysAI4S