Institution profile

BioMap

Industry researchasia · cn
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

MergeDNA: Context-aware Genome Modeling with Dynamic Tokenization through Token Merging

Nov 17, 2025

Genomic sequence modeling faces two fundamental challenges: highly non-uniform information density and the absence of natural minimal lexical units, rendering conventional single-base or static DNA tokenization approaches ill-suited to genomic structural complexity. To address this, we propose a unified framework integrating dynamic tokenization with context-aware pretraining. We introduce differentiable Token Merging—the first such application in genomics—to enable adaptive base-level aggregation. We further design a latent-variable Transformer with hierarchical attention, jointly enforcing local window constraints and global contextual modeling to support selective token identification and reconstruction. Our framework jointly optimizes tokenization and representation learning via two end-to-end objectives: merged-token reconstruction and adaptive masked modeling. Evaluated across three major DNA benchmarks and multi-omics tasks, our method consistently outperforms state-of-the-art tokenization strategies and large-scale DNA foundation models under both fine-tuning and zero-shot settings.

0 citationsRead paper

EnTao-GPM: DNA Foundation Model for Predicting the Germline Pathogenic Mutations

Jul 29, 2025

In precision medicine, accurately distinguishing benign polymorphisms from pathogenic germline variants remains a critical challenge. To address this, we propose a novel framework integrating evolutionary information with interpretable AI: first, cross-species targeted pretraining on multi-organism genomic data leverages evolutionary conservation to improve pathogenicity modeling—especially in noncoding regions; second, task-specific fine-tuning on ClinVar and HGMD couples a DNA foundation model with a large language model (LLM) to jointly perform variant classification and generate statistically grounded, clinically interpretable explanations. Our method achieves significant performance gains over state-of-the-art tools on ClinVar, notably improving accuracy for both SNVs and non-SNV variants—including indels and splice-site alterations. The framework delivers efficient, reliable computational support for automated genetic testing, clinical variant interpretation, and personalized therapeutic intervention.

0 citationsRead paper

TrinityDNA: A Bio-Inspired Foundational Model for Efficient Long-Sequence DNA Modeling

Jul 25, 2025

Traditional sequence models struggle to capture long-range dependencies and biologically relevant structural features in DNA, limiting their effectiveness in gene function prediction and regulatory mechanism inference. To address this, we propose the first biology-informed foundation model for long DNA sequences. Our approach innovatively incorporates Groove Fusion to encode DNA’s 3D groove geometry, gated reverse-complement (GRC) modeling to explicitly represent double-stranded symmetry, and integrates multi-scale attention with an evolutionary training strategy for unified prokaryotic and eukaryotic genome modeling. We concurrently release the first long-sequence DNA benchmark dataset specifically designed for coding sequence (CDS) annotation. On gene function prediction and regulatory element identification tasks, our model significantly outperforms state-of-the-art methods, achieving superior accuracy and generalization across diverse genomic contexts. This work advances the practical deployment of long-sequence genomic foundation models.

0 citationsRead paper

Life-Code: Central Dogma Modeling with Multi-Omics Sequence Unification

Feb 11, 2025

This work addresses the limitation of existing biomolecular pre-trained models, which neglect cross-omics interactions among DNA, RNA, and proteins. We propose the first unified multi-omics modeling framework grounded in the Central Dogma. Methodologically: (1) we introduce a novel nucleotide representation paradigm driven by reverse transcription and reverse translation; (2) we design a codon-aware tokenizer and a hybrid long-sequence Transformer architecture; and (3) we integrate masked modeling pre-training with knowledge distillation from protein language models to enable end-to-end modeling—from coding sequences to tertiary protein structures. Our framework achieves state-of-the-art performance across 12 downstream tasks spanning genomics, transcriptomics, and proteomics. It significantly improves accuracy in cross-omics functional prediction and enhances biological interpretability through mechanistically grounded representations.

0 citationsRead paper
Recent publications

Latest Papers

MergeDNA: Context-aware Genome Modeling with Dynamic Tokenization through Token Merging

Nov 17, 2025

Genomic sequence modeling faces two fundamental challenges: highly non-uniform information density and the absence of natural minimal lexical units, rendering conventional single-base or static DNA tokenization approaches ill-suited to genomic structural complexity. To address this, we propose a unified framework integrating dynamic tokenization with context-aware pretraining. We introduce differentiable Token Merging—the first such application in genomics—to enable adaptive base-level aggregation. We further design a latent-variable Transformer with hierarchical attention, jointly enforcing local window constraints and global contextual modeling to support selective token identification and reconstruction. Our framework jointly optimizes tokenization and representation learning via two end-to-end objectives: merged-token reconstruction and adaptive masked modeling. Evaluated across three major DNA benchmarks and multi-omics tasks, our method consistently outperforms state-of-the-art tokenization strategies and large-scale DNA foundation models under both fine-tuning and zero-shot settings.

0 citationsRead paper

EnTao-GPM: DNA Foundation Model for Predicting the Germline Pathogenic Mutations

Jul 29, 2025

In precision medicine, accurately distinguishing benign polymorphisms from pathogenic germline variants remains a critical challenge. To address this, we propose a novel framework integrating evolutionary information with interpretable AI: first, cross-species targeted pretraining on multi-organism genomic data leverages evolutionary conservation to improve pathogenicity modeling—especially in noncoding regions; second, task-specific fine-tuning on ClinVar and HGMD couples a DNA foundation model with a large language model (LLM) to jointly perform variant classification and generate statistically grounded, clinically interpretable explanations. Our method achieves significant performance gains over state-of-the-art tools on ClinVar, notably improving accuracy for both SNVs and non-SNV variants—including indels and splice-site alterations. The framework delivers efficient, reliable computational support for automated genetic testing, clinical variant interpretation, and personalized therapeutic intervention.

0 citationsRead paper

TrinityDNA: A Bio-Inspired Foundational Model for Efficient Long-Sequence DNA Modeling

Jul 25, 2025

Traditional sequence models struggle to capture long-range dependencies and biologically relevant structural features in DNA, limiting their effectiveness in gene function prediction and regulatory mechanism inference. To address this, we propose the first biology-informed foundation model for long DNA sequences. Our approach innovatively incorporates Groove Fusion to encode DNA’s 3D groove geometry, gated reverse-complement (GRC) modeling to explicitly represent double-stranded symmetry, and integrates multi-scale attention with an evolutionary training strategy for unified prokaryotic and eukaryotic genome modeling. We concurrently release the first long-sequence DNA benchmark dataset specifically designed for coding sequence (CDS) annotation. On gene function prediction and regulatory element identification tasks, our model significantly outperforms state-of-the-art methods, achieving superior accuracy and generalization across diverse genomic contexts. This work advances the practical deployment of long-sequence genomic foundation models.

0 citationsRead paper

Life-Code: Central Dogma Modeling with Multi-Omics Sequence Unification

Feb 11, 2025

This work addresses the limitation of existing biomolecular pre-trained models, which neglect cross-omics interactions among DNA, RNA, and proteins. We propose the first unified multi-omics modeling framework grounded in the Central Dogma. Methodologically: (1) we introduce a novel nucleotide representation paradigm driven by reverse transcription and reverse translation; (2) we design a codon-aware tokenizer and a hybrid long-sequence Transformer architecture; and (3) we integrate masked modeling pre-training with knowledge distillation from protein language models to enable end-to-end modeling—from coding sequences to tertiary protein structures. Our framework achieves state-of-the-art performance across 12 downstream tasks spanning genomics, transcriptomics, and proteomics. It significantly improves accuracy in cross-omics functional prediction and enhances biological interpretability through mechanistically grounded representations.

0 citationsRead paper