Brain2Speech-Net: Intelligible, Real-Time Brain-to-Speech Synthesis Without Text Decoding

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决瘫痪患者因失去语言能力导致的交流障碍,本文提出Brain2Speech-Net,通过直接将神经活动转换成语音而不需文本解码的方法,实现了高可懂度与实时性。
📝 Abstract
The loss of speech limits communication for individuals with paralysis. Restoring speech by synthesizing it directly from neural activity is challenging: intracortical data are scarce and lack aligned targets, so most systems rely on cascaded neural-to-text-to-speech pipelines that add latency and propagate errors. We present Brain2Speech-Net, among the first single-stage frameworks to remain intelligible under limited data while removing intermediate text decoding. A differentiable phoneme bottleneck preserves linguistic structure without explicit text decoding. A lightweight deep-HMM aligner then maps this bottleneck to contextual phoneme representations in a TTS latent space. It learns monotonic alignment between neural recordings and phoneme segments without frame-level supervision, inheriting strong acoustic priors for data-efficient training. On an intracortical dataset, Brain2Speech-Net achieves strong intelligibility in objective and listening tests while running faster than real time. Unlike cascaded systems that incur high latency and direct speech-unit models that lack intelligibility, it delivers both intelligible and real-time speech.
Problem

Research questions and friction points this paper is trying to address.

speech synthesis
neural activity
paralysis
intelligibility
real-time
Innovation

Methods, ideas, or system contributions that make the work stand out.

single-stage framework
differentiable phoneme bottleneck
lightweight deep-HMM aligner
monotonic alignment
real-time speech
🔎 Similar Papers
No similar papers found.