Institution profile

Bose Corporation

Industry researchnorthamerica · us
Official website
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

Real-Time System for Audio-Visual Target Speech Enhancement

Sep 25, 2025

To address the severe degradation of target speech in single-channel audio by environmental noise and interfering speakers—and the limited robustness of conventional audio-only methods—this paper proposes the first interactive audio-visual speech enhancement system capable of real-time operation on commodity CPUs. The system jointly processes raw single-channel audio and lip-motion visual features extracted from a pre-trained audio-visual speech recognition model, employing an end-to-end architecture for speech separation and enhancement. It requires no GPU acceleration, supports synchronous real-time input from microphone and camera, and delivers enhanced speech directly to headphones with minimal latency. Experiments demonstrate consistent performance across diverse realistic noise conditions and multi-talker interference scenarios, exhibiting strong generalization and significantly outperforming audio-only baselines. This work bridges a critical gap by introducing a lightweight, interactive, and CPU-efficient audio-visual speech enhancement solution.

0 citationsRead paper
Recent publications

Latest Papers

Real-Time System for Audio-Visual Target Speech Enhancement

Sep 25, 2025

To address the severe degradation of target speech in single-channel audio by environmental noise and interfering speakers—and the limited robustness of conventional audio-only methods—this paper proposes the first interactive audio-visual speech enhancement system capable of real-time operation on commodity CPUs. The system jointly processes raw single-channel audio and lip-motion visual features extracted from a pre-trained audio-visual speech recognition model, employing an end-to-end architecture for speech separation and enhancement. It requires no GPU acceleration, supports synchronous real-time input from microphone and camera, and delivers enhanced speech directly to headphones with minimal latency. Experiments demonstrate consistent performance across diverse realistic noise conditions and multi-talker interference scenarios, exhibiting strong generalization and significantly outperforming audio-only baselines. This work bridges a critical gap by introducing a lightweight, interactive, and CPU-efficient audio-visual speech enhancement solution.

0 citationsRead paper