ViSAR: Training-Free Adaptive-$k$ Retrieval for Visual Document Question Answering

📅 2026-09-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
ViSAR通过在嵌入空间中构建查询条件页面级相似度矩阵,动态决定检索页面数量,解决了固定top-k检索方法导致的LVLM延迟增加和答案准确性下降问题。
📝 Abstract
Document Visual Question Answering (DocVQA) often leverages Retrieval-Augmented Generation (RAG), where late-interaction encoders are commonly used to identify document pages relevant to a user query, before answer generation by a Large Vision-Language Model (LVLM). Existing approaches typically retrieve a fixed top-$k$ number of pages regardless of query complexity, which increases LVLM latency and may degrade answer accuracy. We introduce ViSAR (Visual Semantic Activation Retrieval), a training-free adaptive-$k$ retrieval method for late-interaction visual document retrieval. ViSAR operates directly in the embedding space to construct a query-conditioned page-level similarity matrix that highlights query-relevant semantics and dynamically determines the number of pages to retrieve. Across multiple encoders and LVLMs, ViSAR retrieves compact, query-adapted page sets that reduce RAG latency by up to 58.7\%, while maintaining or improving answer accuracy compared with fixed top-$k$ and adaptive retrieval heuristics. Furthermore, we show that the similarity matrix structure correlates with answer accuracy, suggesting future directions for retrieval quality-aware document understanding.
Problem

Research questions and friction points this paper is trying to address.

DocVQA
Retrieval-Augmented Generation
Large Vision-Language Model
adaptive retrieval
query complexity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Training-Free
Adaptive-$k$ Retrieval
Similarity Matrix
Late-Interaction
Visual Document Question Answering
🔎 Similar Papers
No similar papers found.
A
Adrien Mialland
INSA Lyon, CNRS, LIRIS UMR 5205, F-69621 Villeurbanne, France
Marc Plantevit
Marc Plantevit
Professor at EPITA, Deputy head of EPITA Research Lab (LRE)
Data miningDatabasePattern MiningMachine LearningxAI
J
Julien Gallois
Lowit, FR-69003 Lyon, France
C
Céline Robardet
INSA Lyon, CNRS, LIRIS UMR 5205, F-69621 Villeurbanne, France