Whence the Voice? Self-supervised Dual-source Audio-Visual Localisation via Selective Convergence
This work addresses the mutual dependency between audio separation and sound source localization in multi-source audio-visual scenarios by proposing a two-stage self-supervised framework. Leveraging a selective convergence mechanism from contrastive learning, the method first localizes the dominant sound source and then iteratively uncovers additional sources using the initial localization as a prior, thereby mimicking human auditory attention. To mitigate existing evaluation biases, the study innovatively incorporates pixel-level semantic segmentation masks to establish a more equitable spatial alignment benchmark. On two-source benchmarks, the proposed approach substantially outperforms all existing self-supervised methods and even surpasses certain weakly supervised techniques on key metrics, advancing the state of the art in label-free multi-source sound localization.