CAVE-NAV: VLM-Based Autonomous 3D Navigation in Underwater Cave Environments

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决水下洞穴环境中自主导航的问题,提出了一种基于视觉-语言模型(VLM)和链式思维推理的框架,利用多模态输入信息推断可航行方向。
📝 Abstract
Autonomous navigation in underwater cave environments is essential for search-and-rescue operations, scientific exploration, and emergency egress. Traditional navigation systems commonly depend on dense visual features for localization and mapping. In underwater caves, however, visual degradation can undermine feature-based localization, sonar-based mapping may yield overly conservative obstacle representations, and communication constraints preclude real-time human guidance. To address these limitations, we propose an autonomous underwater cave navigation framework that leverages a vision-language model (VLM) with Chain-of-Thought (CoT) reasoning to infer navigable directions from environmental cues, including light intensity gradients, passage morphology, and geometric complexity, captured through multimodal inputs comprising RGB imagery, depth maps, and sonar-based vertical-clearance measurements, thereby supporting safe 3D navigation through confined cave passages. High-fidelity simulations across multiple cave topologies demonstrate that the proposed framework completes all evaluated end-to-end traversals without collisions while maintaining safe clearance from cave boundaries.
Problem

Research questions and friction points this paper is trying to address.

autonomous navigation
underwater cave environments
visual degradation
feature-based localization
communication constraints
Innovation

Methods, ideas, or system contributions that make the work stand out.

vision-language model
Chain-of-Thought reasoning
multimodal inputs
💼 Related Jobs
No related jobs found.
Z
Zhenqi Wu
Embodied Robotics and Automation Lab, University of South Florida, Tampa, FL 33620, USA
Y
Yuanjie Lu
Computer Science Department, George Mason University, Fairfax, VA 22032, USA
Y
Yisheng Zhang
Department of Mechanical Engineering, University of Maryland, College Park, MD 20742, USA
Miao Yu
Miao Yu
Professor of Mechanical Engineering, University of Maryland
Metamaterialphotonic sensorsacoustic sensorssensors for robotics
X
Xuesu Xiao
Computer Science Department, George Mason University, Fairfax, VA 22032, USA
J
Jaejeong Shin
Naval Architecture and Ocean Engineering, Seoul National University, Seoul 08826, South Korea
Xiaomin Lin
Xiaomin Lin
Assistant Prof, University of South Florida
AI for goodRobotics for scienceRobotics for good