BoostDream: Efficient Refining for High-Quality Text-to-3D Generation from Multi-View Diffusion
To address the longstanding trade-off between quality and efficiency in text-to-3D generation, this paper proposes a plug-and-play, efficient 3D refinement framework that elevates coarse, feedforward-generated 3D assets to high-fidelity levels within seconds. Methodologically, we introduce the first 3D model distillation mechanism, design a multi-view-aware Score Distillation Sampling (SDS) loss, and incorporate joint guidance from normal maps and text prompts—thereby overcoming the “Janus dilemma” of SDS, where geometric accuracy and rendering speed are conventionally at odds. The framework supports diverse differentiable 3D representations—including NeRF and Gaussian Splatting—without requiring retraining. Extensive experiments demonstrate consistent superiority over state-of-the-art baselines across geometric completeness, texture realism, and inference speed, achieving synergistic improvements in both quality and efficiency.