SSPS: Self-Supervised Positive Sampling for Robust Self-Supervised Speaker Verification
In self-supervised speaker verification, conventional same-utterance positive sampling causes models to over-rely on channel-specific cues and suffer from high intra-speaker variance. To address this, we propose Self-Supervised Positive Sampling (SSPS), which retrieves cross-condition positive samples—i.e., utterances from the same speaker but different recording conditions—within the latent space. SSPS innovatively integrates K-means clustering assignments with a dynamic memory queue to enable label-free, channel-agnostic positive retrieval, thereby decoupling self-supervised learning from recording-condition dependencies. The method is compatible with both SimCLR and DINO frameworks and is optimized via contrastive learning. On VoxCeleb1-O, DINO-SSPS and SimCLR-SSPS achieve EERs of 2.53% and 2.57%, respectively—substantially outperforming prior state-of-the-art methods. Notably, SimCLR-SSPS yields a 58% relative EER reduction.