HybridSB-MoE: Dual-Domain Schrödinger Bridges with Scene-Adaptive Expert Routing for Speech Enhancement

📅 2026-08-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses three key challenges in generative speech enhancement: spectral-domain models disrupting phase information, waveform-domain models neglecting harmonic structure, and the high computational cost and training-inference mismatch of Schrödinger Bridge (SB) methods. To overcome these limitations, the authors propose HybridSB-MoE, a framework that jointly models speech in both spectral and waveform domains. It leverages asymmetric uncertainty fusion, a Top-2 heterogeneous mixture-of-experts (MoE) routing mechanism spanning five distinct architectures, and path consistency regularization to achieve efficient, high-quality enhancement. Theoretical analysis provides an error bound for few-step inference based on the 2-Wasserstein distance. Evaluated on VoiceBank+DEMAND, the method outperforms existing diffusion and SB baselines with fewer inference steps, matching the performance of consistency-distilled few-step approaches.
📝 Abstract
Generative speech enhancement faces three gaps: spectral models capture harmonic structure but often disrupt phase, waveform models preserve phase but miss harmonics, and Schrödinger Bridges (SB) shorten transport from noise to clean speech but leave inference cost only loosely tied to training. We propose HybridSB-MoE, a dual-domain framework that fills these gaps through three contributions unified by a single asymmetric design principle. (i) Asymmetric uncertainty fusion: The spectral path captures epistemic uncertainty via expert disagreement, while the waveform bridge models aleatoric variance through stochastic dynamics. We fuse them asymmetrically, allowing the mixing weight to adapt to distinct error regimes rather than average predictions. (ii) Heterogeneous MoE with top-k=2 routing across five distinct architectural archetypes, where architectural diversity makes the epistemic signal indicate which inductive bias fails rather than small perturbations among similar experts. (iii) Discretization bound (Theorem 1): path-consistency and trajectory regularizers together bound the K-step bridge sampling error in 2-Wasserstein distance at rate K-alpha, making small-K inference an objective-level guarantee rather than an empirical claim. On VoiceBank+DEMAND, HybridSB-MoE outperforms diffusion- and SB-based baselines at their step budgets while remaining competitive with consistency-distilled few-step methods.
Problem

Research questions and friction points this paper is trying to address.

speech enhancement
Schrödinger Bridges
spectral modeling
waveform modeling
generative models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Schrödinger Bridges
Mixture of Experts
Dual-Domain Speech Enhancement
Uncertainty Fusion
Path Consistency
🔎 Similar Papers
Z
Zhengyi Lu
Department of Computer Science and Engineering, Oakland University
A
Aswini Sivakumar
Department of Computer Science and Engineering, Oakland University
J
Jie Hu
Department of Computer Science and Engineering, Oakland University
Yao Qiang
Yao Qiang
Oakland University
Trustworthy AINatural Language ProcessingLarge Language ModelMachine Learning