SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of knowledge transfer caused by the decoupling of spatial perception and reasoning. We propose the first unified framework for native multimodal generation that reformulates 3D reconstruction, correspondence, and reasoning as instruction-following generation tasks. By leveraging sequential output and geometric field generation, our approach enables shared representations and joint optimization across heterogeneous tasks. Incorporating instruction-conditioned generation with spatially supervised learning, the model achieves competitive performance on multiple benchmarks. These results effectively validate the feasibility of handling diverse spatial tasks within a single framework, overcoming the inherent constraints of traditional modular approaches and establishing a new paradigm for integrated spatial intelligence.
📝 Abstract
Spatial perception and reasoning from visual observations require recovering geometric structure, establishing correspondences, and understanding spatial relations. Existing approaches typically address these capabilities separately using task-specific architectures or external geometric modules, limiting knowledge transfer among complementary representations of the same physical scene. We introduce SPARGen, a unified multimodal framework that casts 3D reconstruction, dense correspondence, and spatial reasoning as instruction-conditioned generation tasks. SPARGen serializes compact structured and linguistic outputs as token sequences while generating dense geometric fields in image-aligned forms, enabling spatial supervision to jointly shape shared representations within a native multimodal generative model. Experiments across benchmarks for 3D reconstruction, correspondence, and spatial reasoning show that SPARGen achieves competitive performance across heterogeneous spatial tasks within a single native multimodal generative framework.
Problem

Research questions and friction points this paper is trying to address.

Spatial Perception
Spatial Reasoning
Multimodal Generation
Knowledge Transfer
Innovation

Methods, ideas, or system contributions that make the work stand out.

Native Multimodal Generation
Unified Spatial Framework
Instruction-Conditioned Generation
Joint Representation Learning
Dense Geometric Fields
💼 Related Jobs
No related jobs found.
J
Jinsheng Quan
Zhejiang University
J
Jianhua Li
SenseTime Research
S
Siyi Xie
Peking University
X
Xuanke Shi
SenseTime Research
K
Kewang Deng
SenseTime Research
Z
Zukai Chen
SenseTime Research
Feifei Shao
Feifei Shao
Zhejiang Univiersity
Machine learningcomputer visionweakly supervised learningactive learning
Lei Yang
Lei Yang
SenseTime Research
Machine LearningComputer Vision
Quan Wang
Quan Wang
sensetime
Y
Yawei Luo
Zhejiang University