State-Conditioned Visual Evidence Retrieval for Fine-Grained Perception in Document Vision-Language Models

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为了解决文档解析中细粒度视觉感知问题,提出了一种基于状态条件的视觉证据检索方法(SCVER),并通过空间引导学习目标(SGLO)稳定训练过程。
📝 Abstract
Compared with typical vision-language tasks, document parsing places stronger demands on fine-grained visual perception. Existing vision-language model (VLM)-based parsing approaches rely on globally compressed visual tokens, where fine-grained details are entangled within a single representation and repeatedly accessed during decoding. However, we observe that the visual evidence for each prediction is typically localized and conditioned on the current decoding state, whereas such representations must be accessed in full at every decoding step, resulting in inefficient computation. To address this mismatch, we formulate perception as state-conditioned visual evidence retrieval (SCVER) during autoregressive decoding. The model operates on a compact global representation for coarse structure and retrieves a small set of relevant high-resolution regions conditioned on the current token state. This coarse-to-fine design enables on-demand access to fine-grained visual cues, relieving globally shared representations from encoding all fine-grained details. We further find that learning such state-conditioned retrieval in VLMs is challenging and unstable. To stabilize this process, we introduce a Spatially-Guided Learning Objective (SGLO) to guide the retrieval process. Experiments on document parsing benchmarks show that SCVER improves robustness under reduced input resolution and achieves a better accuracy-efficiency trade-off, demonstrating the effectiveness of on-demand visual evidence retrieval for fine-grained perception.
Problem

Research questions and friction points this paper is trying to address.

visual evidence
fine-grained perception
document parsing
autoregressive decoding
visual tokens
Innovation

Methods, ideas, or system contributions that make the work stand out.

state-conditioned visual evidence retrieval
coarse-to-fine design
Spatially-Guided Learning Objective
fine-grained perception
🔎 Similar Papers
Mingxu Chai
Mingxu Chai
Fudan University
C
Chenyu Liu
College of Computer Science and Artificial Intelligence, Fudan University
Z
Ziyu Shen
College of Computer Science and Artificial Intelligence, Fudan University
Jiazheng Zhang
Jiazheng Zhang
Fudan University
Large Language ModelNatural Language ProcessingData Mining
Kaidi Zhang
Kaidi Zhang
Purdue University
roboticstactile sensing
Ruoyu Chen
Ruoyu Chen
Institute of Information Engineering, Chinese Academy of Sciences.
Explainable AITrustworthy AIFoundation Model
J
Jun Long
ByteDance, LarkAI
J
Jihua Kang
ByteDance, LarkAI
T
Tao Gui
College of Computer Science and Artificial Intelligence, Fudan University
Qi Zhang
Qi Zhang
Fudan University
SAGINsatellite routing