MGFF-TDNN: A Multi-Granularity Feature Fusion TDNN Model with Depth-Wise Separable Module for Speaker Verification

📅 2025-05-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Traditional speaker verification models (e.g., TDNN, ECAPA-TDNN) overly rely on long-term contextual information while neglecting fine-grained speaker characteristics, leading to limited discriminative power. To address this, we propose the Multi-granularity Feature Fusion TDNN (M-TDNN). Our method introduces three key innovations: (1) a 2D depthwise separable convolutional frontend to enhance local time-frequency modeling; (2) the first integration of phoneme-level feature pooling with TDNN, explicitly capturing phoneme-scale discriminative cues; and (3) a triple-domain fusion mechanism combining time-frequency, contextual, and fine-grained features. Evaluated on VoxCeleb1, M-TDNN achieves state-of-the-art performance—significantly reducing EER—while requiring fewer parameters and lower computational cost than standard TDNN and ECAPA-TDNN baselines.

Technology Category

Application Category

📝 Abstract
In speaker verification, traditional models often emphasize modeling long-term contextual features to capture global speaker characteristics. However, this approach can neglect fine-grained voiceprint information, which contains highly discriminative features essential for robust speaker embeddings. This paper introduces a novel model architecture, termed MGFF-TDNN, based on multi-granularity feature fusion. The MGFF-TDNN leverages a two-dimensional depth-wise separable convolution module, enhanced with local feature modeling, as a front-end feature extractor to effectively capture time-frequency domain features. To achieve comprehensive multi-granularity feature fusion, we propose the M-TDNN structure, which integrates global contextual modeling with fine-grained feature extraction by combining time-delay neural networks and phoneme-level feature pooling. Experiments on the VoxCeleb dataset demonstrate that the MGFF-TDNN achieves outstanding performance in speaker verification while remaining efficient in terms of parameters and computational resources.
Problem

Research questions and friction points this paper is trying to address.

Capturing fine-grained voiceprint information in speaker verification
Fusing multi-granularity features for robust speaker embeddings
Balancing performance and efficiency in speaker verification models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-granularity feature fusion for speaker verification
Depth-wise separable convolution for feature extraction
Combines global and fine-grained feature modeling
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Y
Ya Li
College of Computer Science, South-Central Minzu University, Wuhan, China
B
Bin Zhou
College of Computer Science, South-Central Minzu University, Wuhan, China
B
Bo Hu
Wuhan Dongxin Tongbang Information Technology Co., Ltd., Wuhan, China