Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对视觉-语言模型在物理世界空间推理上的局限,提出FactoSR框架,通过分解为平面对应、深度一致性和时间可逆性三个子目标,强化4D一致性,提升3D和4D推理能力。
📝 Abstract
Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain fundamentally ``flat'' when reasoning about the physical world. We argue that this spatial bottleneck stems from a profound dimensional mismatch: while VLMs are trained to interpret 2D projections, true spatial reasoning demands the recovery of latent 3D geometry and temporal continuity. To conquer this high-dimensional complexity, we advocate a shift from monolithic learning to a ``divide and conquer'' paradigm. We present FactoSR, a factorized reinforcement learning framework that explicitly interpret the dimensions collapsed by visual projection. At its core, FactoSR decomposes the monolithic problem of world-consistent reasoning into three orthogonal, geometric sub-objectives: planar correspondence ($XY$), depth consistency ($Z$), and temporal reversibility ($T$). By optimizing these verifiable constraints within a unified policy learning mechanism, we effectively transform an ill-posed projection recovery problem into a series of tangible reasoning steps. Extensive evaluations on multi-view and video benchmarks demonstrate that this elegant decomposition yields substantial gains in 3D and 4D reasoning, achieving a 5.9% boost on VSI-Bench and 4.5% on All-Angles-Bench. Our findings suggest that reinforcing explicit, factorized 4D consistency is a critical step toward evolving VLMs into robust, world-aware reasoners.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
spatial reasoning
3D geometry
temporal continuity
dimensional mismatch
Innovation

Methods, ideas, or system contributions that make the work stand out.

Factorized Reinforcement Learning
4D Consistency
Spatial Reasoning
Dimensional Mismatch