FOMO-3D: Using Vision Foundation Models for Long-Tailed 3D Object Detection
This work addresses the challenge of detecting rare yet safety-critical 3D objects—such as construction workers—under long-tailed distributions in autonomous driving. The authors propose a multimodal two-stage detection framework that, for the first time, leverages vision foundation models (OWLv2 and Metric3Dv2) to provide semantic and depth priors. A novel camera branch is designed to incorporate these priors, and an attention-based mechanism is employed to fuse LiDAR point cloud features with image features, thereby enhancing both proposal generation and refinement. Experiments on real-world driving data demonstrate that the proposed method significantly improves 3D detection performance on long-tailed categories, validating the effectiveness of integrating vision foundation model priors with multimodal fusion strategies.