Online Segment Any 3D Thing as Instance Tracking

๐Ÿ“… 2025-12-08
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Existing online 3D segmentation methods rely on predefined object queries and neglect temporal dynamics in perception, rendering them vulnerable to viewpoint changes, occlusions, and fragmented features from vision foundation models (VFMs), thereby yielding incoherent instance associations and limited holistic understanding. This work pioneers modeling online 3D segmentation as a spatiotemporal instance tracking task. We propose a sparse object query-based framework for temporal feature propagation, integrating long-range instance association with short-term observation refinement, and introduce spatiotemporal consistency learning to mitigate occlusion and feature fragmentation. Our method operates directly on VFM-driven 3D point clouds without requiring additional annotations. On ScanNet200, it achieves a +2.8 AP gain over ESAM; consistent improvements are observed across ScanNet, SceneNN, and 3RScan. The approach significantly enhances fine-grained, temporally coherent instance perception in dynamic environments.

Technology Category

Application Category

๐Ÿ“ Abstract
Online, real-time, and fine-grained 3D segmentation constitutes a fundamental capability for embodied intelligent agents to perceive and comprehend their operational environments. Recent advancements employ predefined object queries to aggregate semantic information from Vision Foundation Models (VFMs) outputs that are lifted into 3D point clouds, facilitating spatial information propagation through inter-query interactions. Nevertheless, perception is an inherently dynamic process, rendering temporal understanding a critical yet overlooked dimension within these prevailing query-based pipelines. Therefore, to further unlock the temporal environmental perception capabilities of embodied agents, our work reconceptualizes online 3D segmentation as an instance tracking problem (AutoSeg3D). Our core strategy involves utilizing object queries for temporal information propagation, where long-term instance association promotes the coherence of features and object identities, while short-term instance update enriches instant observations. Given that viewpoint variations in embodied robotics often lead to partial object visibility across frames, this mechanism aids the model in developing a holistic object understanding beyond incomplete instantaneous views. Furthermore, we introduce spatial consistency learning to mitigate the fragmentation problem inherent in VFMs, yielding more comprehensive instance information for enhancing the efficacy of both long-term and short-term temporal learning. The temporal information exchange and consistency learning facilitated by these sparse object queries not only enhance spatial comprehension but also circumvent the computational burden associated with dense temporal point cloud interactions. Our method establishes a new state-of-the-art, surpassing ESAM by 2.8 AP on ScanNet200 and delivering consistent gains on ScanNet, SceneNN, and 3RScan datasets.
Problem

Research questions and friction points this paper is trying to address.

Addresses dynamic 3D segmentation as instance tracking for embodied agents.
Enhances temporal coherence via long-term association and short-term update mechanisms.
Mitigates fragmentation from Vision Foundation Models with spatial consistency learning.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reconceptualizes 3D segmentation as instance tracking
Uses object queries for temporal information propagation
Introduces spatial consistency learning to mitigate fragmentation
๐Ÿ”Ž Similar Papers
H
Hanshi Wang
State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), CASIA
Z
Zijian Cai
AutoLab, School of Artificial Intelligence, Shanghai Jiao Tong University
J
Jin Gao
State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), CASIA
Y
Yiwei Zhang
State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), CASIA
W
Weiming Hu
State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), CASIA
K
Ke Wang
KargoBot
Zhipeng Zhang
Zhipeng Zhang
School of Artificial Intelligence, Shanghai Jiao Tong University
Computer Vision๏ผŒObject Tracking and Segmentation