SincQDR-VAD: A Noise-Robust Voice Activity Detection Framework Leveraging Learnable Filters and Ranking-Aware Optimization
To address the poor robustness of voice activity detection (VAD) under noisy and resource-constrained conditions, and the misalignment between conventional classification losses and evaluation metrics such as AUROC, this paper proposes a compact, efficient end-to-end VAD framework. Methodologically: (i) a learnable Sinc bandpass filter is employed to construct a noise-robust spectral frontend, enhancing feature discriminability; (ii) a novel Quadratic Difference Ranking Loss is introduced to explicitly optimize the relative ranking of speech versus non-speech frames, thereby directly maximizing AUROC. Experiments on multiple benchmark datasets demonstrate consistent improvements—AUROC increases by 1.2–2.8% and F2-score by 3.5–5.1%—while the model requires only 69% of the parameters of current state-of-the-art methods. The proposed approach thus achieves superior accuracy, low inference latency, and high parameter efficiency.