UniFusion: Sparse-View 4D Reconstruction via Unified Spatio-temporal Depth Alignment

📅 2026-09-05
📈 Citations: 0
Influential: 0
📄 PDF
📝 Abstract
In this paper, we address the challenging problem of 4D reconstruction from sparse-view videos. This setup usually relies on monocular depth estimation to provide priors for the reconstruction model. A key challenge arises from limited cross-view overlap and temporal variation, making monocular depth predictions inconsistent across views and time. Existing methods align spatial and temporal dimensions in separate stages, requiring foreground segmentation masks while failing to leverage temporal cues for cross-view alignment. Contrary to these methods, we propose a unified spatial-temporal depth alignment framework that jointly resolves cross-view and cross-time inconsistencies without distinguishing foreground/background. Our method represents depth maps across views and time as a set of spatio-temporal neural fields. This representation not only yields fast convergence, but also captures spatio-temporal correlation among depth maps implicitly, without dependence on external segmentation/tracking models. We also propose a multi-view depth-order loss while leveraging the classic scale-and-shift-invariant loss to further improve the final depth quality. The aligned depths initialize and supervise Gaussian splatting models for 4D reconstruction. Experiments on Ego-Exo4D and EgoHuman demonstrate that our improved depth alignment substantially benefits dynamic Gaussian-splatting-based reconstruction methods for novel-time/view synthesis and geometry accuracy/consistency.
Problem

Research questions and friction points this paper is trying to address.

4D reconstruction
sparse-view videos
monocular depth estimation
cross-view overlap
temporal variation
Innovation

Methods, ideas, or system contributions that make the work stand out.

unified spatio-temporal depth alignment
spatio-temporal neural fields
multi-view depth-order loss
🔎 Similar Papers
Y
Yongzhe Lyu
State Key Laboratory of General Artificial Intelligence, BIGAI, Beijing, China
S
Shaofei Wang
State Key Laboratory of General Artificial Intelligence, BIGAI, Beijing, China
Y
Yixin Chen
State Key Laboratory of General Artificial Intelligence, BIGAI, Beijing, China
Siyuan Huang
Siyuan Huang
Beijing Institute for General Artificial Intelligence (BIGAI)
Embodied AI3D VisionRobotics3D Scene Understanding