IsoNet: Spatially-aware audio-visual target speech extraction in complex acoustic environments

πŸ“… 2026-05-14
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the challenge of target speech extraction on compact devices equipped with small-aperture microphone arrays, where conventional beamforming and monaural neural models struggle to perform effectively. To overcome this limitation, we propose IsoNetβ€”a multimodal U-Net-based mask estimation network tailored for four-microphone arrays. IsoNet uniquely integrates complex-valued multi-channel STFT features, GCC-PHAT spatial cues, face-conditioned visual embeddings, and DOA-assisted supervision, enabling user-selectable target speaker extraction. Trained with three curriculum learning strategies, the model significantly outperforms traditional approaches across a wide SNR range from βˆ’1 to 10 dB, achieving a SI-SDR of 9.31 dB (a 4.85 dB improvement), PESQ of 2.13, and STOI of 0.84, thereby effectively mitigating performance degradation in low-SNR and small-aperture scenarios.
πŸ“ Abstract
Target speech extraction remains difficult for compact devices because monaural neural models lack spatial evidence and classical beamformers lose resolving power when the microphone aperture is only a few centimetres. We present IsoNet, a user-selectable audio-visual target speech extraction system for a compact 4-microphone array. IsoNet combines complex multi-channel STFT features, GCC-PHAT spatial cues, face-conditioned visual embeddings, and auxiliary direction-of-arrival supervision inside a U-Net mask estimation network. Three curriculum variants were trained on 25,000 simulated VoxCeleb mixtures with progressively difficult SNR regimes. On a hard test set spanning -1 to 10 dB SNR, IsoNet-CL1 achieves 9.31 dB SI-SDR, a 4.85 dB improvement over the mixture, with PESQ 2.13 and STOI 0.84. Oracle delay-and-sum and MVDR beamformers degrade the same mixtures by 4.82 dB and 6.08 dB SI-SDRi, respectively, showing that the proposed learned multimodal conditioning solves a regime where conventional spatial filtering is ineffective. Ablation studies show consistent gains from visual conditioning, GCC-PHAT features, and extended delay-bin encoding. The results establish a compact-array, face-selectable speech extraction baseline under controlled simulation and identify the remaining barriers to real deployment, especially phase reconstruction, multi-interferer mixtures, and simulation-to-real transfer.
Problem

Research questions and friction points this paper is trying to address.

target speech extraction
compact microphone array
complex acoustic environments
spatial cues
audio-visual separation
Innovation

Methods, ideas, or system contributions that make the work stand out.

audio-visual speech extraction
compact microphone array
spatial cues
multimodal conditioning
U-Net mask estimation
πŸ”Ž Similar Papers
πŸ’Ό Related Jobs
No related jobs found.
D
Dinanath Pathya
Department of Electronics and Computer Engineering, Thapathali Campus, Institute of Engineering, Tribhuvan University, Kathmandu, Nepal
S
Sajen Maharjan
Department of Electronics and Computer Engineering, Thapathali Campus, Institute of Engineering, Tribhuvan University, Kathmandu, Nepal
B
Binita Adhikari
Department of Electronics and Computer Engineering, Thapathali Campus, Institute of Engineering, Tribhuvan University, Kathmandu, Nepal
I
Ishwor Raj Pokharel
Department of Electronics and Computer Engineering, Thapathali Campus, Institute of Engineering, Tribhuvan University, Kathmandu, Nepal