CVSBench: A Comprehensive Benchmark for Cross-view Spatial Reasoning and Dreaming

๐Ÿ“… 2026-06-21
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study investigates whether vision-language models possess human-like spatial reasoning capabilities in cross-view scenarios, such as between satellite and street-level imagery. To this end, the authors introduce a large-scale multitask benchmark comprising 3,297 cross-view image pairs, 9,468 object annotations, and 40,679 question-answer pairs, supporting tasks including cross-view visual question answering, object localization, and view recognition. They propose a novel visual-spatial representation method that integrates 3D scene imagination, substantially improving model performance in understanding object and layout consistency under extreme viewpoint variations. Experimental results demonstrate that purely language-based reasoning yields limited effectiveness, whereas incorporating visual-spatial imagination significantly enhances the modelโ€™s spatial cognition abilities.
๐Ÿ“ Abstract
Humans can effortlessly reason about scenes across different viewpoints, yet it remains unclear whether Vision-Language Models (VLMs) possess similar cross-view spatial abilities. Satellite-street scene pairs, with their complex contexts and extreme viewpoint variations, provide an ideal testbed. Motivated by this, we introduce CVSBench, a large-scale benchmark for evaluating cross-view spatial reasoning through satellite-street pairs. This benchmark supports multiple tasks, including cross-view VQA, cross-view grounding, and viewpoint identification. CVSBench comprises 3,297 cross-view image groups with 9,468 object-level annotations and 40,679 question-answer (QA) pairs, enabling systematic and controlled evaluation of cross-view spatial reasoning. Extensive evaluations reveal that advanced VLMs struggle to maintain object-level and layout consistency under drastic viewpoint changes. To bridge this gap towards human-like spatial cognition, we investigate two categories of approaches: spatially grounded reasoning and the incorporation of cognitive map inputs. Our findings demonstrate that language-only reasoning yields marginal improvements, while incorporating visual spatial imagination via a 3D scene imagination pipeline substantially improves cross-view reasoning. These results highlight the necessity of explicit visual-spatial representations for robust spatial cognition in VLMs. Our data and code are released at https://huggingface.co/datasets/zlyzlyzly/CVSBench.
Problem

Research questions and friction points this paper is trying to address.

cross-view spatial reasoning
vision-language models
satellite-street scene pairs
spatial cognition
viewpoint variation
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-view spatial reasoning
vision-language models
3D scene imagination
spatial grounding
cognitive map
๐Ÿ”Ž Similar Papers
Ruixun Liu
Ruixun Liu
Undergraduates of Xi'an Jiaotong University
computer vision
L
Lingyu Zhang
Faculty of Electronic and Information Engineering, Xiโ€™an Jiaotong University
L
Lanxuan Xue
Faculty of Electronic and Information Engineering, Xiโ€™an Jiaotong University
Kaiyu Li
Kaiyu Li
Wilfrid Laurier University, Canada
Data governance and Data preparationData market and Data economy
Bowen Fu
Bowen Fu
Department of Automation, Tsinghua University
3D VisionPose EstimationRobotic Manipulation
X
Xiangyong Cao
Faculty of Electronic and Information Engineering, Xiโ€™an Jiaotong University; Ministry of Education Key Laboratory of Intelligent Networks and Network Security