OpenBelief-Nav: Evidence-Preserving Object Memory for Open-Vocabulary Language-Guided Navigation

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the loss of task hypotheses caused by premature semantic fixation in open-vocabulary navigation by proposing an Evidence-Retaining Object Memory mechanism. By preserving observation-level phrases and reliability cues, this method constructs vocabulary-agnostic belief representations that decouple geometric-visual features and facilitate multi-candidate verification, thereby enabling flexible retrieval and dynamic correction. Experimental results demonstrate significant mIoU improvements on ScanNet200 and Replica datasets. Furthermore, the model achieves success rates of 60/78 on HM3D-YCB and 8/10 in real-world target confirmation tasks. These findings confirm that the proposed approach effectively enhances both the generalization capability and robustness of open-vocabulary navigation systems.
📝 Abstract
Open-vocabulary 3D scene graphs provide compact semantic memory for language-guided navigation, but mapped objects are often exposed through a single fused feature or committed semantic label. Such commitment can remove minority yet task-relevant hypotheses from the task-time interface. We present OpenBelief-Nav, an evidence-preserving object memory that retains observation-level phrases, reliability cues, and frame-mask provenance while maintaining separate aggregate geometric and visual representations. Semantically related phrases are consolidated into a vocabulary-independent object belief from which task-specific readouts perform fixed-vocabulary projection or free-form retrieval. On five ScanNet200 and eight Replica scenes, full-belief projection achieves mIoU scores of 0.2742 and 0.2912, compared with 0.2393 and 0.2701 for a matched early-commit readout. Across 78 HM3D-YCB navigation trials, consensus and early-commit retrieval each achieve 60/78 successes, compared with 58/78 for belief-weighted retrieval and 55/78 for DualMap. Across 20 Unitree G1 runs organized as 10 matched evaluation cases, a correction policy permitting at most two verified candidate attempts improves target-confirmation success from 6/10 to 8/10 relative to top-1-only execution. Code will be released upon acceptance at https://openbelief-nav.github.io/.
Problem

Research questions and friction points this paper is trying to address.

Open-vocabulary navigation
3D scene graphs
Early semantic commitment
Object memory
Language-guided navigation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evidence-Preserving Object Memory
Open-Vocabulary 3D Scene Graphs
Vocabulary-Independent Object Belief
Language-Guided Navigation
Correction Policy
🔎 Similar Papers
No similar papers found.
D
Dinh Tuan Nguyen
VinMotion, Inc., Vietnam
Anh Dao
Anh Dao
Undergraduate Student, Michigan State University
Vision-languageMultimodal LLMEmbodied AILLM
P
Phuong Nam Dang
VinMotion, Inc., Vietnam
Quan-Dung Pham
Quan-Dung Pham
VinMotion, Inc., Vietnam
T
Tuyen P. Le
VinMotion, Inc., Vietnam
T
Truong Nguyen
VinMotion, Inc., Vietnam
Quan Nguyen
Quan Nguyen
University of Southern California
ControlRoboticsOptimization