Institution profile

Shanghai Ocean University

Academic institutionasia · cn
Official website
Research library6linked papers
Opportunities0open roles
Selected work

Representative Papers

An Empirical Study on the Robustness of YOLO Models for Underwater Object Detection

Sep 22, 2025

Underwater object detection faces severe challenges including low-level feature degradation (e.g., texture, edge, and color distortion), noise interference, and class imbalance due to optical distortions inherent in aquatic environments. Method: This work conducts a systematic robustness evaluation of YOLOv8–v12 across six simulated underwater conditions using the DUO and Roboflow100 datasets (10,000 annotated images), employing cross-model and cross-environment benchmarking. It further proposes a noise-aware sample injection strategy and enhancement-domain fine-tuning to improve generalization under noise perturbations and real underwater domains. Contribution/Results: We identify—for the first time—the robustness bottlenecks of YOLO models underwater: although YOLOv12 achieves the highest overall accuracy, it exhibits extreme sensitivity to noise; detection performance is predominantly constrained by sample quantity and instance frequency. Our lightweight training paradigm and targeted image enhancement significantly boost robustness and domain adaptability, empirically validating their efficacy for underwater detection.

0 citationsRead paper

A Structured Review of Underwater Object Detection Challenges and Solutions: From Traditional to Large Vision Language Models

Sep 10, 2025

Underwater object detection (UOD) faces five core challenges: severe image degradation, small and highly deformable object scales, scarcity of annotated data, stringent real-time inference requirements, and poor generalization of existing models. To address these, this work systematically analyzes the challenges and surveys methodological evolution. Crucially, it pioneers the integration of large vision-language models (LVLMs) into UOD: leveraging DALL·E 3 to generate high-fidelity synthetic underwater imagery, and employing Florence-2 for multimodal fine-tuning and cross-domain transfer. Experiments demonstrate substantial improvements in detection accuracy and robustness under complex underwater conditions—particularly in realistic scene modeling—while highlighting persistent bottlenecks in small-object localization and dynamic scene adaptation. This study bridges a critical gap by establishing the first LVLM-based framework for UOD, offering a novel pathway toward data-efficient learning and lightweight deployment.

0 citationsRead paper
Recent publications

Latest Papers

An Empirical Study on the Robustness of YOLO Models for Underwater Object Detection

Sep 22, 2025

Underwater object detection faces severe challenges including low-level feature degradation (e.g., texture, edge, and color distortion), noise interference, and class imbalance due to optical distortions inherent in aquatic environments. Method: This work conducts a systematic robustness evaluation of YOLOv8–v12 across six simulated underwater conditions using the DUO and Roboflow100 datasets (10,000 annotated images), employing cross-model and cross-environment benchmarking. It further proposes a noise-aware sample injection strategy and enhancement-domain fine-tuning to improve generalization under noise perturbations and real underwater domains. Contribution/Results: We identify—for the first time—the robustness bottlenecks of YOLO models underwater: although YOLOv12 achieves the highest overall accuracy, it exhibits extreme sensitivity to noise; detection performance is predominantly constrained by sample quantity and instance frequency. Our lightweight training paradigm and targeted image enhancement significantly boost robustness and domain adaptability, empirically validating their efficacy for underwater detection.

0 citationsRead paper

A Structured Review of Underwater Object Detection Challenges and Solutions: From Traditional to Large Vision Language Models

Sep 10, 2025

Underwater object detection (UOD) faces five core challenges: severe image degradation, small and highly deformable object scales, scarcity of annotated data, stringent real-time inference requirements, and poor generalization of existing models. To address these, this work systematically analyzes the challenges and surveys methodological evolution. Crucially, it pioneers the integration of large vision-language models (LVLMs) into UOD: leveraging DALL·E 3 to generate high-fidelity synthetic underwater imagery, and employing Florence-2 for multimodal fine-tuning and cross-domain transfer. Experiments demonstrate substantial improvements in detection accuracy and robustness under complex underwater conditions—particularly in realistic scene modeling—while highlighting persistent bottlenecks in small-object localization and dynamic scene adaptation. This study bridges a critical gap by establishing the first LVLM-based framework for UOD, offering a novel pathway toward data-efficient learning and lightweight deployment.

0 citationsRead paper