🤖 AI Summary
This study addresses the limitation of knowledge transfer caused by the decoupling of spatial perception and reasoning. We propose the first unified framework for native multimodal generation that reformulates 3D reconstruction, correspondence, and reasoning as instruction-following generation tasks. By leveraging sequential output and geometric field generation, our approach enables shared representations and joint optimization across heterogeneous tasks. Incorporating instruction-conditioned generation with spatially supervised learning, the model achieves competitive performance on multiple benchmarks. These results effectively validate the feasibility of handling diverse spatial tasks within a single framework, overcoming the inherent constraints of traditional modular approaches and establishing a new paradigm for integrated spatial intelligence.
📝 Abstract
Spatial perception and reasoning from visual observations require recovering geometric structure, establishing correspondences, and understanding spatial relations. Existing approaches typically address these capabilities separately using task-specific architectures or external geometric modules, limiting knowledge transfer among complementary representations of the same physical scene. We introduce SPARGen, a unified multimodal framework that casts 3D reconstruction, dense correspondence, and spatial reasoning as instruction-conditioned generation tasks. SPARGen serializes compact structured and linguistic outputs as token sequences while generating dense geometric fields in image-aligned forms, enabling spatial supervision to jointly shape shared representations within a native multimodal generative model. Experiments across benchmarks for 3D reconstruction, correspondence, and spatial reasoning show that SPARGen achieves competitive performance across heterogeneous spatial tasks within a single native multimodal generative framework.