Institution profile

Inception Institute of Artificial Intelligence

Academic institutionasia · ae
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Prot42: a Novel Family of Protein Language Models for Target-aware Protein Binder Generation

Apr 06, 2025

Designing high-affinity protein binders remains challenging when target 3D structures and binding sites are unknown. To address this, we propose the first purely sequence-driven, target-aware de novo protein binder generation framework. Our method introduces an autoregressive protein language model (pLM) decoder architecture supporting ultra-long sequences of up to 8,192 residues. Trained on large-scale unlabeled protein sequences, it implicitly integrates evolutionary, structural, and functional information—requiring no target structural input. Crucially, it enables target-sequence-conditioned generation, breaking from conventional structure-dependent paradigms. On high-affinity protein binder design and sequence-specific DNA-binding protein generation, our approach significantly outperforms structure-based baselines including AlphaProteo. The code and models are publicly released to facilitate reproducible, accessible protein engineering.

0 citationsRead paper

Gene42: Long-Range Genomic Foundation Model With Dense Attention

Mar 20, 2025

This work addresses the challenge of modeling long-range dependencies in genomic sequences by introducing the first genome foundation model capable of single-nucleotide-resolution modeling over sequences up to 192 kbp. Methodologically, it adopts a LLaMA-style decoder-only architecture with dense self-attention—eschewing convolutional or state-space modules that impose locality constraints—and employs a phased context expansion strategy to ensure stable training on ultra-long sequences. Key contributions include: (i) the first empirical validation in genomics that dense attention scales effectively to hundred-kilobase sequences, establishing a new paradigm for long-range dependency modeling; and (ii) state-of-the-art performance across diverse tasks—including biological sequence classification, regulatory region identification, chromatin accessibility prediction, pathogenicity assessment of genetic variants, and cross-species classification—while achieving both low perplexity and high sequence reconstruction fidelity. The model is publicly released on Hugging Face.

0 citationsRead paper

Chem42: a Family of chemical Language Models for Target-aware Ligand Generation

Mar 20, 2025

Current chemical language models (cLMs) struggle to effectively incorporate target-specific information, limiting their utility in target-aware de novo ligand generation. To address this, we introduce the first target-aware generative cLM family, which achieves deep integration of structural target priors into cLMs via cross-modal collaboration with the protein language model Prot42—enabling atom-level protein–ligand interaction modeling for the first time. Our approach comprises three key components: multimodal representation learning, joint fine-tuning with Prot42, and target-conditioned molecular decoding. Evaluated across multiple protein targets, the model substantially improves chemical validity (>98%), target selectivity, and binding affinity prediction accuracy, while dramatically narrowing the search space for viable candidates. The open-source models establish new state-of-the-art performance on the Hugging Face chemical benchmark, offering a novel paradigm for synthesizable, highly specific ligand design.

0 citationsRead paper
Recent publications

Latest Papers

Prot42: a Novel Family of Protein Language Models for Target-aware Protein Binder Generation

Apr 06, 2025

Designing high-affinity protein binders remains challenging when target 3D structures and binding sites are unknown. To address this, we propose the first purely sequence-driven, target-aware de novo protein binder generation framework. Our method introduces an autoregressive protein language model (pLM) decoder architecture supporting ultra-long sequences of up to 8,192 residues. Trained on large-scale unlabeled protein sequences, it implicitly integrates evolutionary, structural, and functional information—requiring no target structural input. Crucially, it enables target-sequence-conditioned generation, breaking from conventional structure-dependent paradigms. On high-affinity protein binder design and sequence-specific DNA-binding protein generation, our approach significantly outperforms structure-based baselines including AlphaProteo. The code and models are publicly released to facilitate reproducible, accessible protein engineering.

0 citationsRead paper

Gene42: Long-Range Genomic Foundation Model With Dense Attention

Mar 20, 2025

This work addresses the challenge of modeling long-range dependencies in genomic sequences by introducing the first genome foundation model capable of single-nucleotide-resolution modeling over sequences up to 192 kbp. Methodologically, it adopts a LLaMA-style decoder-only architecture with dense self-attention—eschewing convolutional or state-space modules that impose locality constraints—and employs a phased context expansion strategy to ensure stable training on ultra-long sequences. Key contributions include: (i) the first empirical validation in genomics that dense attention scales effectively to hundred-kilobase sequences, establishing a new paradigm for long-range dependency modeling; and (ii) state-of-the-art performance across diverse tasks—including biological sequence classification, regulatory region identification, chromatin accessibility prediction, pathogenicity assessment of genetic variants, and cross-species classification—while achieving both low perplexity and high sequence reconstruction fidelity. The model is publicly released on Hugging Face.

0 citationsRead paper

Chem42: a Family of chemical Language Models for Target-aware Ligand Generation

Mar 20, 2025

Current chemical language models (cLMs) struggle to effectively incorporate target-specific information, limiting their utility in target-aware de novo ligand generation. To address this, we introduce the first target-aware generative cLM family, which achieves deep integration of structural target priors into cLMs via cross-modal collaboration with the protein language model Prot42—enabling atom-level protein–ligand interaction modeling for the first time. Our approach comprises three key components: multimodal representation learning, joint fine-tuning with Prot42, and target-conditioned molecular decoding. Evaluated across multiple protein targets, the model substantially improves chemical validity (>98%), target selectivity, and binding affinity prediction accuracy, while dramatically narrowing the search space for viable candidates. The open-source models establish new state-of-the-art performance on the Hugging Face chemical benchmark, offering a novel paradigm for synthesizable, highly specific ligand design.

0 citationsRead paper