GLaQ: Grounding Latent Queries in Visual Evidence for Multimodal Reasoning

📅 2026-08-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the loss of fine-grained visual evidence during text generation in multimodal reasoning by proposing the GLaQ framework. Replacing autoregressive latent reasoning with grounded latent queries, this method directly anchors original visual tokens and integrates local view supervision with task-level reinforcement learning to enable precise recovery and reuse of visual evidence without external operations. Experimental results across five benchmarks demonstrate that GLaQ outperforms baselines by 5.99%–9.66%, significantly surpassing existing visual latent approaches. These findings confirm that GLaQ effectively mitigates visual information degradation in multimodal large language model reasoning, offering a robust solution for preserving critical visual details throughout the generative process.
📝 Abstract
Chain-of-thought reasoning has substantially improved the problem-solving capabilities of multimodal large language models. Fine-grained visual evidence, however, remains difficult to preserve and reuse across text-based reasoning steps. To address this limitation, tool-augmented thinking-with-images methods maintain visual access externally by revisiting or manipulating the image, but require predefined tools and additional inference-time processing. As an internal alternative, continuous visual latent reasoning retains intermediate computation in hidden states. However, its prevailing autoregressive construction makes each latent state depend on its predecessors, so later states may repeat information already present in the latent sequence rather than capture complementary visual details. We introduce GLaQ, a grounded latent-query framework that replaces sequential latent rollout with a fixed set of context-conditioned queries grounded in the original visual tokens. The grounded queries are reinjected for answer generation, providing direct and coordinated access to source visual evidence. We train GLaQ with localized-view supervision followed by reinforcement learning under task-level rewards. Across five benchmarks for fine-grained visual understanding and perception, GLaQ-7B gains 5.99--9.66\% over its base model and leads all compared visual latent methods, suggesting that direct query-to-image grounding can recover localized evidence from the full image without external visual operations or autoregressive latent rollouts.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Reasoning
Visual Evidence
Latent Reasoning
Fine-grained Visual Understanding
Autoregressive Latent States
Innovation

Methods, ideas, or system contributions that make the work stand out.

Grounded Latent Queries
Visual Evidence Grounding
Non-autoregressive Latent Reasoning
Reinforcement Learning with Task Rewards
Multimodal Chain-of-Thought
🔎 Similar Papers
No similar papers found.
Z
Zesheng Yang
Xi'an Jiaotong University
Lingling Zhang
Lingling Zhang
Assistant Professor, Xi'an Jiaotong University
Computer visionFew-shot learningZero-shot learning
X
Xinyu Zhang
Xi'an Jiaotong University
C
Cheng Zhang
Xi'an Jiaotong University
P
Pengyu Li
Xi'an Jiaotong University
H
Heng Wang
Xi'an Jiaotong University
L
Lin Wu
Xi'an Jiaotong University