Efficient4D: Fast Dynamic 3D Object Generation from a Single-view Video
To address the challenges of missing 4D annotations and low end-to-end optimization efficiency in monocular video-based dynamic 3D reconstruction, this work proposes a two-stage decoupled paradigm: first generating multi-view temporally consistent images via a diffusion model, then driving 4D Gaussian Splatting for explicit reconstruction. We introduce an inconsistency-aware confidence-weighted loss and a lightweight Score Distillation Sampling (SDS) loss, significantly improving robustness under sparse-view conditions. Compared to Consistent4D, our method accelerates training tenfold (10 minutes vs. 120 minutes), enables real-time continuous trajectory rendering, and achieves state-of-the-art novel-view synthesis quality. To the best of our knowledge, this is the first work to organically integrate generative modeling with explicit 4D reconstruction, establishing a new paradigm for efficient, high-fidelity dynamic scene reconstruction.