Dynamic-Robust Photometric-Semantic Reconstruction for Open-Vocabulary 3D Scene Understanding

📅 2026-08-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决动态环境中空间特征错位问题,提出SPAR架构隔离瞬态噪声,并通过端到端训练结合运动估计与多视图学习,实现高质量新视角合成和语义理解。
📝 Abstract
The integration of novel view synthesis (NVS) and open-vocabulary segmentation (OVS) has recently yielded powerful feed-forward 3D foundation models. However, their inherent reliance on static-scene assumptions leads to severe misalignment of spatial features in unconstrained dynamic environments. To bridge this critical gap, we propose SPAR, a novel joint semantic-geometric encoding architecture that explicitly isolates transient dynamic noise prior to latent space aggregation. Furthermore, we introduce a dynamic-region-aware end-to-end training paradigm that structurally couples motion estimation with multi-view visual and semantic learning. This unified approach enables the network to inherently resolve motion conflicts and distill multi-view consistent, temporally stable scene representations from dynamic inputs. Extensive experiments on the challenging D-RE10K benchmark demonstrate that SPAR achieves state-of-the-art performance. Our end-to-end approach achieves exceptional novel view synthesis quality, yielding a PSNR of 22.15 dB and 23.33 dB given only 3 and 4 input views respectively. Despite being trained in a self-supervised manner, our model achieves an mIoU of 88.5% for motion mask prediction. Furthermore, our analysis reveals a strong inter-task synergy between photometric scene reconstruction and semantic understanding, where semantic synthesis learning consistently enhances photometric fidelity in novel view rendering. Code will be available at https://github.com/dmucby/SPAR.
Problem

Research questions and friction points this paper is trying to address.

novel view synthesis
open-vocabulary segmentation
dynamic environments
spatial feature misalignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

dynamic-region-aware
joint semantic-geometric encoding
end-to-end training paradigm
motion estimation
multi-view consistent
💼 Related Jobs
No related jobs found.
B
Boyu Cai
School of Information Science and Technology, ShanghaiTech University; State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), Beijing Key Laboratory of Super Intelligent Security of Multi-Modal Information, Institute of Automation, Chinese Academy of Sciences (CASIA)
Li Yang
Li Yang
Insititute of Software, Chinese Academy of Sciences
Software EngineeringArtificial Intelligence
Y
Yan Xu
Electronic Engineering Department, The Chinese University of Hong Kong
W
Wei Liu
State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), Beijing Key Laboratory of Super Intelligent Security of Multi-Modal Information, Institute of Automation, Chinese Academy of Sciences (CASIA)
Nian Liu
Nian Liu
Mohamed bin Zayed University of Artificial Intelligence
Computer VisionSaliencyFew-shot Learning3D PerceptionMultimodal AI
S
Sikui Zhang
State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), Beijing Key Laboratory of Super Intelligent Security of Multi-Modal Information, Institute of Automation, Chinese Academy of Sciences (CASIA)
Y
Yan Wang
Deepeleph Intelligent Technology
Chunfeng Yuan
Chunfeng Yuan
National Laboratory of Pattern Recognition, Institute of Automation, Chinese Academy of Sciences
computer visionPattern RecognitionMachine LearningHuman Action RecognitionSparse Representation
W
Weiming Hu
School of Information Science and Technology, ShanghaiTech University; State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), Beijing Key Laboratory of Super Intelligent Security of Multi-Modal Information, Institute of Automation, Chinese Academy of Sciences (CASIA)