MAD: Multi-Alignment MEG-to-Text Decoding
Current non-invasive BCI-based language decoding faces three key bottlenecks: underutilization of magnetoencephalography (MEG) signals, poor cross-sentence generalization, and absence of multimodal fusion. Method: We propose the first end-to-end, multi-aligned MEG-to-text framework for natural language reconstruction from entirely unseen sentences. Our approach introduces a Transformer-based architecture that jointly aligns neural time series, phonemes, and semantics, integrating self-supervised pretraining with cross-modal contrastive learning to systematically unify speech, semantic, and dynamic temporal information. Results: On the Gwilliams dataset, our method achieves a BLEU-1 score of 10.44—improving by 4.95 (+93%) over the strongest baseline—demonstrating substantially enhanced open-vocabulary text generation capability. This work breaks critical limitations in generalizability and multimodal integration for non-invasive brain–computer interface–based language reconstruction.