VCP-DCN: Beyond Visual Concealed Property via Depth Collaborative Network for Camouflaged Object Detection

πŸ“… 2026-07-30
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the limitation of existing camouflaged object detection methods that overlook the modality-specific characteristics of camouflaged objects in depth data, thereby hindering effective multimodal collaboration. To overcome this, we propose a Depth-Collaborative Network (VCP-DCN) that disentangles prototype embeddings to separately model modality-shared and modality-specific features. The network integrates a multimodal dual-attention mechanism to enhance cross-modal local responses and introduces a depth-adaptive injection module for efficient feature fusion. Extensive experiments on three benchmark RGB-D camouflaged object detection datasets demonstrate that our method significantly outperforms state-of-the-art approaches, confirming its effectiveness and robustness.
πŸ“ Abstract
Camouflaged Object Detection (COD) aims to identify and segment camouflaged objects in complex environments, which are often concealed because their color and texture are similar to the background. Several existing COD methods introduce depth maps to boost detection performance via learning complementary RGB-D features, ignoring modality-specific characteristics of concealed objects in the depth domain. To address this issue, we propose a depth collaborative network, called VCP-DCN, to mine distinguishable multi-modality features beyond visual concealed prototype in depth domain. Specifically, VCP-DCN progressively performs multi-modality alignment, interaction, and fusion for the COD task. In the \textbf{alignment} stage, we propose a Separable Prototype Embedding (SPE) module to learn modality-consistency and modality-specific RGB/depth prototype tokens through prototype contrastive learning. Furthermore, we develop a Multi-modality Dual Attention (MDA) module to enhance the cross-modal feature representation through local response maps between modality-consistency RGB/depth prototype tokens and visual tokens on the \textbf{interaction} stage. Finally, we design a Depth Adaptive Injection (DAI) module to adaptively measure contribution of RGB/depth features with a decision-making mechanism, which calculates similarity distance between RGB/depth modality-specific prototype tokens and modality-consistency ones on the \textbf{fusion} stage. Extensive experiments demonstrate the effectiveness of our VCP-DCN on three authoritative datasets.
Problem

Research questions and friction points this paper is trying to address.

Camouflaged Object Detection
depth modality
modality-specific characteristics
RGB-D features
visual concealed property
Innovation

Methods, ideas, or system contributions that make the work stand out.

Camouflaged Object Detection
Depth Collaborative Network
Prototype Contrastive Learning
Multi-modality Fusion
Modality-specific Features
πŸ’Ό Related Jobs
No related jobs found.
S
Songsong Duan
the State Key Laboratory of Integrated Services Networks, School of Telecommunications Engineering, Xidian University, Xi’an, China
Xi Yang
Xi Yang
Xidian University
Computer VisionMachine LearningPattern Recognition
Nannan Wang
Nannan Wang
Professor, Xidian University
Computer VisionMachine LearningPattern Recognition