Exploring Diffusion Transformers for Cross-Modal Augmentation in Multimodal Brain State Decoding

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出CoMA-DiT模型,通过跨模态注意力机制增强多模态脑状态解码训练数据,实验显示该方法在准确性和宏观F1分数上优于现有基线。
📝 Abstract
Multimodal brain state decoding has largely focused on fusing paired modalities for prediction, but has rarely explored how their correspondence can be further exploited to enrich training data and improve multimodal representation learning. To address this gap, we propose CoMA-DiT, a bidirectional cross-modal Diffusion Transformer for latent augmentation that treats paired modalities as sources of mutual generative supervision rather than merely as inputs to be fused. CoMA-DiT conditions velocity prediction on the paired modality through cross-modal attention and adaptively injects the resulting variation via a reliability-gated residual mechanism. Experiments on multimodal auditory attention decoding and emotion recognition showed that CoMA-DiT consistently outperformed 20 representative baselines, achieving absolute gains of 4.28% and 6.70% in accuracy and macro-F1 over the no-augmentation baseline, respectively. Extensive ablation, sensitivity, visualization, and interpretability analyses further demonstrated its robustness, generalizability, and ability to capture functionally relevant cross-modal interactions. These findings support a broader view of multimodal learning: Paired modalities can serve not only as inputs for fusion but also as supervision sources that augment one another.
Problem

Research questions and friction points this paper is trying to address.

multimodal brain state decoding
cross-modal augmentation
representation learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-Modal Augmentation
Diffusion Transformer
Bidirectional
Reliability-Gated Residual Mechanism
Mutual Generative Supervision
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Ziwei Wang
Ziwei Wang
Huazhong University of Science and Technology
Brain-Computer InterfaceDeep Learning
Xingyi He
Xingyi He
Zhejiang University
Computer Vision
Hongbin Wang
Hongbin Wang
Texas A&M University Health Science Center
Biomedical InformaticsCognitive ScienceAI
T
Tianwang Jia
School of Artificial Intelligence and Automation, Huazhong University of Science and Technology, China
B
Bohan Fang
School of Artificial Intelligence and Automation, Huazhong University of Science and Technology, China
D
Dongrui Wu
School of Artificial Intelligence and Automation, Huazhong University of Science and Technology, China