Model as Loss: A Self-Consistent Training Paradigm
Traditional speech enhancement methods rely on handcrafted or pretrained feature-based losses, limiting their ability to model fine-grained signal characteristics. To address this, we propose a “model-as-loss” self-consistency training paradigm: the encoder of an end-to-end differentiable encoder-decoder network serves as a dynamic, task-aware loss function, constructing a self-supervised objective in a discriminative feature space that enforces intrinsic consistency between enhanced outputs and clean speech. Crucially, this loss is fully internal—requiring no external pretrained models (e.g., WavLM or wav2vec)—and emerges solely from the current network’s own representation. Experiments demonstrate that our approach surpasses state-of-the-art deep feature losses based on WavLM/wav2vec on standard benchmarks, yields significant improvements in subjective speech quality (e.g., PESQ, STOI, and MOS), and exhibits superior generalization both within-domain and cross-domain.