BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and Language

πŸ“… 2026-06-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Existing brain signal encoding and decoding methods typically treat tasks in isolation, relying on unimodal alignment and external priors while neglecting the brain’s inherent multimodal integration. This work proposes BrainJanus, the first unified framework that jointly models brain activity, vision, and language. It introduces a unified brain tokenizer to discretize continuous neural activity into tokens, aligns all modalities within a shared Omni representation space, and employs an all-in-one autoregressive architecture enabling arbitrary-to-arbitrary generation. Without task-specific designs, BrainJanus achieves zero-shot generalization while preserving interpretable neuroanatomical topology. Experiments demonstrate that BrainJanus outperforms existing approaches across multiple benchmarks, offering both strong cross-modal generative capabilities and significant potential for neuroscience interpretation.
πŸ“ Abstract
Modeling the bidirectional correspondence between external sensory stimuli and internal neural activity has emerged as a critical frontier in neuroscience. However, existing approaches predominantly treat brain encoding and decoding as isolated tasks, relying heavily on unimodal alignment and external priors while overlooking the brain's intrinsic nature as a multimodal integration system. To address these limitations, we propose BrainJanus, the first unified brain model that integrates brain, vision, and language within a single framework. Specifically, we introduce a Unified Brain Tokenizer to quantize continuous neural dynamics into discrete tokens aligned with visual and linguistic representations in a shared Omni space. Building on this, we utilize an All-in-One autoregressive architecture that leverages next-token prediction to enable seamless any-to-any generation, which encompasses image-to-brain and text-to-brain encoding, and brain-to-image and brain-to-text decoding. Extensive experiments demonstrate that BrainJanus achieves superior performance across diverse benchmarks. Furthermore, our framework exhibits zero-shot generalization and preserves interpretable biological topography, highlighting its potential as a general-purpose brain modeling paradigm. The code is available at \href{https://github.com/HaitaoWuTJU/BrainJanus}{GitHub}.
Problem

Research questions and friction points this paper is trying to address.

brain encoding
brain decoding
multimodal integration
neural activity
sensory stimuli
Innovation

Methods, ideas, or system contributions that make the work stand out.

unified brain model
multimodal integration
brain tokenizer
any-to-any generation
zero-shot generalization
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Haitao Wu
Haitao Wu
Microsoft
NetworkingDatacenterQoSTCP/IPWireless
Q
Qirui Zhang
School of Artificial Intelligence, Tianjin University, Tianjin, China
Z
Zhouheng Yao
Shanghai Artificial Intelligence Laboratory, Shanghai, China
Shangquan Sun
Shangquan Sun
University of Chinese Academy of Sciences
Computer VisionMachine Learning
Qihao Zheng
Qihao Zheng
Shanghai AI Lab
NeuroscienceNeuroAIAI4NeuroAI4Science
M
Mianxin Liu
Shanghai Artificial Intelligence Laboratory, Shanghai, China
C
Chi Zhang
Shanghai Artificial Intelligence Laboratory, Shanghai, China
W
Wanli Ouyang
Shanghai Artificial Intelligence Laboratory, Shanghai, China; The Chinese University of Hong Kong, Hong Kong, China
Chunfeng Song
Chunfeng Song
Shanghai AI Lab
Computer VisionPattern RecognitionAI4Science
Changqing Zhang
Changqing Zhang
Professor, Tianjin University
Machine LearningMultimodal LearningLLM
Jiamin Wu
Jiamin Wu
The Chinese University of Hong Kong, Shanghai AI Lab
Computer VisionFew-Shot learningAI4Science