ShapeGaussian: High-Fidelity 4D Human Reconstruction in Monocular Videos via Vision Priors
This work addresses the distortions and artifacts commonly observed in existing 4D human reconstruction methods from monocular video, which often stem from reliance on parametric templates such as SMPL or high sensitivity to pose estimation errors. To overcome these limitations, we propose the first template-free, high-fidelity 4D reconstruction framework. Our approach leverages 2D visual priors together with a pretrained data-driven model to generate an initial deformable geometry, which is subsequently refined through a neural deformation field and a multi-reference-frame strategy to capture fine dynamic details. By eliminating template constraints entirely, our method effectively mitigates issues caused by occluded keypoints and inaccurate pose estimates. Extensive experiments demonstrate that our approach significantly outperforms state-of-the-art template-based methods in reconstruction accuracy, visual quality, and robustness across diverse everyday motions captured in monocular videos.