Whence the Voice? Self-supervised Dual-source Audio-Visual Localisation via Selective Convergence

📅 2026-08-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the mutual dependency between audio separation and sound source localization in multi-source audio-visual scenarios by proposing a two-stage self-supervised framework. Leveraging a selective convergence mechanism from contrastive learning, the method first localizes the dominant sound source and then iteratively uncovers additional sources using the initial localization as a prior, thereby mimicking human auditory attention. To mitigate existing evaluation biases, the study innovatively incorporates pixel-level semantic segmentation masks to establish a more equitable spatial alignment benchmark. On two-source benchmarks, the proposed approach substantially outperforms all existing self-supervised methods and even surpasses certain weakly supervised techniques on key metrics, advancing the state of the art in label-free multi-source sound localization.
📝 Abstract
Localising multiple sound sources in visual scenes remains a fundamental challenge in multimodal perception due to an inherent circular dependency: separating mixed audio requires knowing source locations, while identifying sound-producing regions requires separated audio signals. In this paper, we focus on the dual-source setting and discover a selective convergence in self-supervised audio-visual learning: when presented with multiple sound sources, contrastive models naturally converge to the most salient audio-visual correspondence rather than attempting to represent all sources equally. This emergent phenomenon, analogous to human selective auditory attention, enables us to break the above circular dependency through a progressive two-stage framework: first, leveraging selective convergence to identify dominant sources, and then exploiting these learned priors to uncover remaining sources. Our self-supervised approach achieves the best performance among self-supervised methods on dual-source benchmarks without requiring any manual annotations, and even surpasses some weakly-supervised approaches \red{on certain metrics. Furthermore, we identify a fundamental evaluation inconsistency in existing benchmarks: comparing continuous localisation heatmaps against bounding-box annotations creates systematic biases, particularly for non-axis-aligned objects where the bounding box includes substantial background regions. To address this, we introduce pixel-level segmentation masks to the existing benchmark, enabling spatially-aligned evaluation. Together, these results suggest that embracing rather than suppressing selectivity offers a scalable, annotation-free route to multi-source localisation.
Problem

Research questions and friction points this paper is trying to address.

audio-visual localisation
multi-source separation
circular dependency
self-supervised learning
selective attention
Innovation

Methods, ideas, or system contributions that make the work stand out.

selective convergence
self-supervised learning
audio-visual localisation
multi-source separation
pixel-level evaluation
💼 Related Jobs
No related jobs found.
H
Han Hu
The MIx Group, School of Computer Science, University of Birmingham, UK
D
Dongheng Lin
The MIx Group, School of Computer Science, University of Birmingham, UK
Yuqi Hou
Yuqi Hou
University of Birmingham
Computer VisionGaze FollowingMultimodal
H
Haotian Li
The MIx Group, School of Computer Science, University of Birmingham, UK
Hyung Jin Chang
Hyung Jin Chang
Associate Professor, University of Birmingham
Computer VisionRoboticsDeep LearningMachine Learning
Jianbo Jiao
Jianbo Jiao
University of Birmingham | University of Oxford
Computer VisionMachine Learning