SSPS: Self-Supervised Positive Sampling for Robust Self-Supervised Speaker Verification

📅 2025-05-20
📈 Citations: 1
Influential: 0
📄 PDF
🤖 AI Summary
In self-supervised speaker verification, conventional same-utterance positive sampling causes models to over-rely on channel-specific cues and suffer from high intra-speaker variance. To address this, we propose Self-Supervised Positive Sampling (SSPS), which retrieves cross-condition positive samples—i.e., utterances from the same speaker but different recording conditions—within the latent space. SSPS innovatively integrates K-means clustering assignments with a dynamic memory queue to enable label-free, channel-agnostic positive retrieval, thereby decoupling self-supervised learning from recording-condition dependencies. The method is compatible with both SimCLR and DINO frameworks and is optimized via contrastive learning. On VoxCeleb1-O, DINO-SSPS and SimCLR-SSPS achieve EERs of 2.53% and 2.57%, respectively—substantially outperforming prior state-of-the-art methods. Notably, SimCLR-SSPS yields a 58% relative EER reduction.

Technology Category

Application Category

📝 Abstract
Self-Supervised Learning (SSL) has led to considerable progress in Speaker Verification (SV). The standard framework uses same-utterance positive sampling and data-augmentation to generate anchor-positive pairs of the same speaker. This is a major limitation, as this strategy primarily encodes channel information from the recording condition, shared by the anchor and positive. We propose a new positive sampling technique to address this bottleneck: Self-Supervised Positive Sampling (SSPS). For a given anchor, SSPS aims to find an appropriate positive, i.e., of the same speaker identity but a different recording condition, in the latent space using clustering assignments and a memory queue of positive embeddings. SSPS improves SV performance for both SimCLR and DINO, reaching 2.57% and 2.53% EER, outperforming SOTA SSL methods on VoxCeleb1-O. In particular, SimCLR-SSPS achieves a 58% EER reduction by lowering intra-speaker variance, providing comparable performance to DINO-SSPS.
Problem

Research questions and friction points this paper is trying to address.

Limitation of same-utterance positive sampling in SSL for speaker verification
Need to encode speaker identity without shared recording conditions
Proposing SSPS to find same-speaker positives with different recording conditions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Supervised Positive Sampling (SSPS) technique
Clustering assignments for positive selection
Memory queue for storing positive embeddings
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
T
Theo Lepage
EPITA Research Laboratory (LRE), France
R
Réda Dehak
EPITA Research Laboratory (LRE), France