Selective Transfer Learning of Cross-Modality Distillation for Monocular 3D Object Detection
Monocular 3D object detection is inherently ill-posed due to the absence of precise depth information, and existing cross-modal knowledge distillation approaches often suffer from negative transfer caused by the modality gap between images and LiDAR. To address this issue, this work proposes MonoSTL, which presents the first systematic analysis of negative transfer in cross-modal distillation and introduces two novel components: Depth-Aware Selective Feature Distillation (DASFD) and Depth-Aware Selective Relation Distillation (DASRD). These modules leverage depth uncertainty to guide positive knowledge transfer and effectively integrate LiDAR-derived depth cues through structural alignment and selective distillation mechanisms. Extensive experiments demonstrate that MonoSTL significantly boosts the performance of various baseline models on both KITTI and NuScenes benchmarks, achieving state-of-the-art results and confirming its effectiveness and generalizability.