FusionBERT: Multi-View Image-3D Retrieval via Cross-Attention Visual Fusion and Normal-Aware 3D Encoder
This work addresses the limitations of existing image-to-3D multimodal retrieval methods, which predominantly rely on single-view images and struggle to handle the multi-view observations typical of real-world scenarios. To overcome this, the authors propose FusionBERT, a novel framework that first introduces a multi-view visual aggregator leveraging cross-attention mechanisms to adaptively fuse complementary features from multiple viewpoints. Additionally, a normal-aware 3D encoder is incorporated to jointly encode point coordinates and surface normals, thereby enhancing geometric representation for models lacking texture or suffering from color degradation. Experimental results demonstrate that the proposed method significantly outperforms current state-of-the-art approaches under both single-view and multi-view settings, establishing a strong baseline for image-to-3D cross-modal retrieval.