Chain of Spatial Thoughts: Modality-Agnostic Spatial Grounding for Vision Language Models

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing vision-language models often rely on additional encoders or architectural modifications for spatial reasoning, incurring substantial computational overhead. This work proposes Space Tokens, a lightweight and architecture-agnostic framework that introduces continuous spatial tokens to implicitly encode 3D scene geometry and object-centric spatial properties through knowledge distillation, seamlessly integrating them into chain-of-thought reasoning. Notably, the approach requires no modification to the backbone model, offering high efficiency, interpretability, and multimodal extensibility. Evaluated on VSI-Bench, it achieves performance gains of 4.3% and 1.3% for Qwen3-VL-8B and SenseNova-SI-1.3, respectively, and sets new state-of-the-art results in object size estimation (79.2%) and room size estimation (75.7%).
📝 Abstract
Spatial understanding is fundamental to embodied intelligence, underpinning applications such as robotic manipulation, embodied navigation, and autonomous driving. Although recent vision-language models (VLMs) have achieved impressive performance on spatial reasoning benchmarks, state-of-the-art approaches typically rely on additional spatial encoders or architectural modifications during inference, increasing computational cost. We introduce Space Tokens, a lightweight, architecture-agnostic framework that equips VLMs with explicit continuous spatial representations without requiring additional inference-time modules. By distilling scene-level 3D geometry and object-centric spatial attributes into continuous latent tokens, our method enables these modalities to be directly incorporated into a chain-of-thought reasoning process, thereby improving the VLM's spatial reasoning capabilities. At the same time, the learned representations can be explicitly decoded to verify that they encode meaningful geometric information, while the unified token interface remains extensible to additional modalities. Experiments on VSI-Bench improve Qwen3-VL-8B by 4.3% and SenseNova-SI-1.3 by 1.3%, while achieving state-of-the-art performance on object size (79.2%) and room size estimation (75.7%). These results demonstrate that continuous spatial tokens provide an effective, interpretable, and computationally efficient mechanism for integrating geometric reasoning into large vision-language models.
Problem

Research questions and friction points this paper is trying to address.

spatial reasoning
vision-language models
spatial grounding
computational efficiency
modality-agnostic
Innovation

Methods, ideas, or system contributions that make the work stand out.

Space Tokens
spatial grounding
vision-language models
chain-of-thought reasoning
modality-agnostic
🔎 Similar Papers
No similar papers found.
H
Hunter Schofield
York University
M
Mohammed Elmahgiubi
Huawei Technologies Canada
M
Mohammad Mahdavian
Huawei Technologies Canada
R
Richard Shi
University of Toronto
J
Jinjun Shan
York University
Amir Rasouli
Amir Rasouli
Noah's Ark Laboratory
RoboticsComputer VisionAutonomous DrivingVisual Attention
Dongfeng Bai
Dongfeng Bai
Huawei Technologies Co., Ltd.
Computer VisionAutonomous DrivingNeural Rendering