NemoSplat: Feed-Forward 4D Gaussian Splatting for Media-Aware Underwater Reconstruction

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决水下环境光散射和动态物体干扰问题,NemoSplat提出了一种前馈4D高斯点绘框架,通过媒体感知的高斯预测器和可提示的动态解缠器实现精准重建。
📝 Abstract
Reconstructing photorealistic scenes in unconstrained underwater environments remains challenging due to severe media-induced light scattering and unpredictable dynamic objects. Recent feed-forward visual foundation models have demonstrated remarkable capabilities in generalized novel view synthesis and tracking. However, when directly applied to aquatic videos, optical attenuation and motion interference fatally corrupt their feature aggregation, leading to severe tracking and reconstruction failures. To overcome these limitations, we present NemoSplat, the first feed-forward 4D Gaussian Splatting framework tailored for media-aware dynamic reconstruction directly from uncalibrated marine videos. Beyond providing robust estimations of camera poses and dense scene depth, we devise a Promptable Dynamic Disentangler that utilizes a confidence-aware fusion strategy of learned dynamic probabilities and optional semantic text priors, effectively isolating massive transient entities. Furthermore, to counteract visual degradation, a Media-Aware Gaussian Predictor is formulated to jointly estimate intrinsic 3D Gaussian attributes alongside physical media parameters, rendering pristine scene appearance in a single forward pass. Additionally, we introduce a large-scale underwater dataset with massive dynamic elements to facilitate training and evaluation. Extensive experiments on our dataset demonstrate that NemoSplat achieves state-of-the-art tracking accuracy and high-fidelity rendering.
Problem

Research questions and friction points this paper is trying to address.

underwater reconstruction
light scattering
dynamic objects
optical attenuation
motion interference
Innovation

Methods, ideas, or system contributions that make the work stand out.

feed-forward 4D Gaussian Splatting
Promptable Dynamic Disentangler
Media-Aware Gaussian Predictor
X
Xiaopeng Guo
The Hong Kong University of Science and Technology, China
W
Wai Chung Tse
The Hong Kong University of Science and Technology, China
Y
Yipeng Zhu
The Hong Kong University of Science and Technology, China
H
Hanwen Zhang
The Hong Kong University of Science and Technology, China
H
Huajian Huang
Beijing Institute of Technology, China
Sai-Kit Yeung
Sai-Kit Yeung
Integrative Systems and Design, Hong Kong University of Science and Technology
Computer VisionComputer GraphicsComputational Design