When Depth Hurts: Reliability-Aware Geometry Distillation for Depth-Free RGB-D Salient Object Detection

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种无需深度图的RGB-D显著物体检测方法,通过从冻结的Depth Anything V2模型中蒸馏几何信息,并结合像素级可靠性估计器来提升RGB-only模型性能。
📝 Abstract
Depth can resolve appearance ambiguity in RGB-D salient object detection (SOD), yet sensor depth is not uniformly reliable. Missing regions, blurred boundaries, and structural artifacts can propagate through multimodal fusion and make an RGB-D detector less accurate than its RGB-only counterpart. Existing quality-aware approaches regulate observed depth but remain dependent on the same potentially defective modality. We propose \method, a reliability-aware geometry distillation framework developed for RGB-D SOD benchmarks without using dataset-provided depth during training or inference. A frozen Depth Anything V2 model serves only as a training-time teacher, transferring dense relative geometry, hierarchical spatial attention, and boundary structure to a compact edge-aware geometry branch. Pooled bidirectional interaction aligns geometry with appearance, and a pixel-wise reliability estimator selectively injects geometry that is compatible with the current RGB representation. The teacher is removed after training, leaving an RGB-only inference network. Trained on 2,985 RGB-mask pairs, \method{} achieves the best or tied-best result in 26 of 36 metric-dataset comparisons against ten recent RGB-D SOD methods, including a 13.4\% relative MAE reduction on ReDWeb-S. When retrained on DUTS-TR, it also improves the strongest prior $F$-measure by 4.2\% on PASCAL-S, showing that the distilled geometry transfers beyond a particular sensor or dataset domain. Code will be released upon publication.
Problem

Research questions and friction points this paper is trying to address.

Depth
RGB-D SOD
Reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

reliability-aware geometry distillation
dense relative geometry
hierarchical spatial attention
pixel-wise reliability estimator
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Xuehao Wang
Xuehao Wang
Zhejiang University
Multi-Task LearningSegment Anything ModelPEFTLLM
J
Jiaxin Hua
University of International Business and Economics
R
Runmei Li
University of International Business and Economics
Zhenyu Wu
Zhenyu Wu
PhD student, Xi'an Jiaotong University
natural language processing
C
Chenglizhao Chen
China University of Petroleum
K
Ke Gu
Beijing University Of Technology
A
Aimin Hao
State Key Laboratory of Virtual Reality Technology and Systems