OVIP-SG: Open-Vocabulary Instance-Preserving Scene Graphs for Mapping and Retrieval of Small, Fine-Grained Objects

📅 2026-08-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
OVIP-SG通过视觉-语言模型和对称3D IoU关联等方法解决了开放词汇感知中实例一致性问题,提高了小、细粒度物体的检测与检索准确性。
📝 Abstract
Integrating open-vocabulary perception into object-level 3D scene graphs is a double-edged sword. While vision-language detectors recover long-tail categories and small, fine-grained objects overlooked by closed-set models, they also tend to fragment large surfaces and merge small objects into larger neighboring objects, compromising instance-level consistency and undermining mapping fidelity. Moreover, existing methods struggle to retrieve previously unmapped targets or determine whether a queried object is absent, hindering robust embodied open-world navigation and exploration. We present OVIP-SG, a unified framework for instance-preserving semantic mapping, functional scene partitioning, and language-guided small, fine-grained object retrieval. OVIP-SG uses a vision-language model (VLM) to enumerate scene-specific categories for robust open-world detection. Symmetric 3D Intersection over Union (IoU) association and area-weighted feature fusion preserve small independent instances, while VLM-inferred object functions partition scenes into compact functional search regions. A four-stage cascaded retrieval pipeline further incorporates voxel voting and determines target absence from exploration coverage. Under a unified evaluation protocol on Replica, OVIP-SG outperforms ConceptGraphs by 6.31 points in class-mean accuracy (mAcc) and 5.15 points in frequency-weighted mIoU (F-mIoU) while achieving a class-agnostic native-instance Panoptic Quality (PQ) of 0.398. It reduces the search area to 21.8% of the indoor floor space and reaches 0.773 balanced accuracy for object-presence classification. Real-world robotic experiments further demonstrate its practical effectiveness.
Problem

Research questions and friction points this paper is trying to address.

open-vocabulary perception
3D scene graphs
instance-level consistency
object retrieval
embodied open-world navigation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Open-Vocabulary Perception
Instance-Preserving
Fine-Grained Object Retrieval
Vision-Language Model
Functional Scene Partitioning
🔎 Similar Papers
T
Tianjing Hao
Xi’an Jiaotong University
Haiyu Lan
Haiyu Lan
University of Calgary
PerceptionLocalization and Mapping for Autonomous VehicleRobotics.
A
Angsong Li
Zhiyuan Innovation (Shanghai) Technology Co., Ltd.
C
Cheng Chen
Zhiyuan Innovation (Shanghai) Technology Co., Ltd.
E
Enyu Li
Zhiyuan Innovation (Shanghai) Technology Co., Ltd.
J
Jiarui Yang
Zhiyuan Innovation (Shanghai) Technology Co., Ltd.
Y
Yuning Su
Zhiyuan Innovation (Shanghai) Technology Co., Ltd.
P
Peiwen Lin
Zhiyuan Innovation (Shanghai) Technology Co., Ltd.
W
Wang Chuang
Zhiyuan Innovation (Shanghai) Technology Co., Ltd.