Grounding Isn't Knowing: Do VLMs Need Object Localization for Spatial Reasoning?

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过多种工具分析了VLMs在空间推理中是否需要精确对象定位,发现粗略的目标参考锚点即可,不需精确定位。
📝 Abstract
Vision-language models (VLMs) can answer spatial questions, yet the mechanisms connecting object grounding to spatial reasoning remain poorly understood. It is underexplored whether spatial reasoning internally requires precise objects localization, or can bypass explicit localization through global layout cues. In this work, we investigate two representative model families, LLaVA-1.5 and Qwen2.5-VL, using a suite of mechanistic interpretability tools, including token ablation, layer-wise probing, attention knockout, and causal mediation analysis. We find that spatial relation prediction follows a staged grounding-to-reasoning process in which object-aligned tokens establish coarse target-reference anchors, while precise bounding-box boundaries are not required. Positional information becomes decodable before relation decisions emerge, and a small set of attention heads mediates the causal effects of both localization and spatial reasoning. The two tasks share early grounding-related processing but ultimately rely on partially distinct specialized pathways. Through rigorous experiments, we provide a token-, layer-, and head-level account of how VLMs transform object grounding into spatial relations, showing that knowing where objects are is not equivalent to knowing how they relate.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Spatial Reasoning
Object Localization
Innovation

Methods, ideas, or system contributions that make the work stand out.

object-aligned tokens
coarse target-reference anchors
global layout cues
attention heads
mechanistic interpretability
🔎 Similar Papers
No similar papers found.
X
Xiwei Liu
Mohamed bin Zayed University of Artificial Intelligence
Y
Yulong Li
Mohamed bin Zayed University of Artificial Intelligence
X
Xinlin Zhuang
Mohamed bin Zayed University of Artificial Intelligence, The Chinese University of Hong Kong
X
Xuhui Li
Mohamed bin Zayed University of Artificial Intelligence
Z
Zhixiang Lu
University Of Liverpool
Haolin Yang
Haolin Yang
University of Chicago
large language modelsnatural language processing
Imran Razzak
Imran Razzak
MBZUAI, Abu Dhabi
Human-Centered AIMedical Image AnalysisMedical Artificial IntelligenceComputational Biology
Yutong Xie
Yutong Xie
Assistant Professor, Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)
Medical image analysisComputer visionDeep learningMulti-modal learning