SonarLLM: A Native Sonar--Optical Multimodal Large Language Model for Underwater Perception

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决水下感知问题,提出SonarLLM模型,结合声纳和光学模态,通过特定编码器、特征增强及层次融合方法,提高在不同能见度下的感知可靠性。
📝 Abstract
Reliable underwater perception requires complementary sensing under variable visibility. Optical cameras capture appearance and semantics but degrade rapidly with turbidity, whereas imaging sonar preserves geometry while exhibiting distinct range-azimuth structure and acoustic artifacts. Existing MLLMs, built primarily on optical encoders, are therefore ill-suited to model sonar or adaptively exploit sonar-optical complementarity. We propose SonarLLM, a sonar-optical MLLM that treats sonar as a native perceptual modality. It combines a sonar-specific encoder, modality-specific physics-aware feature enhancement, and reliability-aware hierarchical fusion to align acoustic structure with optical semantics and dynamically adjust their contributions as sensing quality changes. We also introduce SonarBench, a paired benchmark that spans four tasks: recognition, counting, visual question answering, and captioning; and, across the benchmark, three input settings: sonar-only, optical-only, and fusion. By fixing the scene and sonar observation while varying optical degradation, SonarBench enables controlled measurement of cross-modal complementarity. SonarLLM achieves 72.0% macro accuracy across sonar-only recognition, counting, and VQA, outperforming the strongest baseline by 34.4 percentage points, and 68.7% under fusion, exceeding the best baseline by 25.1 points. For recognition and counting, the fusion-over-optical gain grows from 6.0 to 36.0 points as turbidity increases, indicating the increasing complementary value of sonar under controlled optical degradation. Together, these results show that robust heterogeneous perception depends not only on adding sonar, but on representing and weighting it according to its sensing characteristics.
Problem

Research questions and friction points this paper is trying to address.

underwater perception
complementary sensing
optical cameras
imaging sonar
multimodal large language model
Innovation

Methods, ideas, or system contributions that make the work stand out.

sonar-optical multimodal
physics-aware feature enhancement
reliability-aware hierarchical fusion
cross-modal complementarity
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Cong Su
Cong Su
Yale University
atomic engineeringelectron microscopy2D materialsquantum emitters
L
Longxuan Ma
Faculty of Information Engineering and Automation, Kunming University of Science and Technology; Yunnan Key Laboratory of Artificial Intelligence, Kunming, China
L
Ling Dong
Faculty of Information Engineering and Automation, Kunming University of Science and Technology; Yunnan Key Laboratory of Artificial Intelligence, Kunming, China
G
Guofeng Tang
Faculty of Information Engineering and Automation, Kunming University of Science and Technology; Yunnan Key Laboratory of Artificial Intelligence, Kunming, China
Weijie Yin
Weijie Yin
ByteDance
Vision Language ModelDeep LearningAI4S
H
Haohui Chen
Faculty of Information Engineering and Automation, Kunming University of Science and Technology; Yunnan Key Laboratory of Artificial Intelligence, Kunming, China
Zhengtao Yu
Zhengtao Yu
Kunming University of Science and Technology