MUSE: Model-based Uncertainty-aware Similarity Estimation for zero-shot 2D Object Detection and Segmentation
To address the challenges of localizing unseen categories and achieving robust cross-modal matching in zero-shot 2D object detection and segmentation, this paper proposes MUSE—a training-free, model-driven framework. MUSE leverages multi-view 2D renderings of 3D unseen objects as templates and performs cross-modal matching against candidate regions extracted from query images. It introduces a novel joint similarity metric—integrating both absolute and relative similarities—and incorporates uncertainty-aware object priors. Furthermore, it employs class-token embedding fusion and generalized mean pooling (GeM) to calibrate candidate region reliability. Evaluated on the BOP Challenge 2025, MUSE achieves state-of-the-art performance in zero-shot detection and segmentation, securing first place across all tracks: Classic Core, H3, and Industrial.