A Phylogenetic Approach to Genomic Language Modeling

📅 2025-03-04

📈 Citations: 0

✨ Influential: 0

career value

188K/year

🤖 AI Summary

Current genomic language models (gLMs) exhibit limited performance in identifying evolutionarily constrained elements in mammalian genomes. To address this, we propose PhyloGPN—a self-supervised genomic language model that explicitly integrates phylogenetic modeling. PhyloGPN is the first gLM to incorporate multi-species whole-genome alignments directly into its loss function, enabling evolutionary signal utilization during training while requiring only a single input sequence—without alignment—at inference time. Furthermore, it constrains nucleotide substitution dynamics using a phylogenetic tree, enhancing biological interpretability. Experiments demonstrate that PhyloGPN significantly outperforms baseline models in predicting functionally disruptive variants and exhibits strong cross-species generalization. By jointly ensuring phylogenetic rigor and practical deployability, PhyloGPN establishes a novel paradigm for interpretable, high-accuracy functional annotation of genomes.

Technology Category

Application Category

📝 Abstract

Genomic language models (gLMs) have shown mostly modest success in identifying evolutionarily constrained elements in mammalian genomes. To address this issue, we introduce a novel framework for training gLMs that explicitly models nucleotide evolution on phylogenetic trees using multispecies whole-genome alignments. Our approach integrates an alignment into the loss function during training but does not require it for making predictions, thereby enhancing the model's applicability. We applied this framework to train PhyloGPN, a model that excels at predicting functionally disruptive variants from a single sequence alone and demonstrates strong transfer learning capabilities.

Problem

Research questions and friction points this paper is trying to address.

Improving genomic language models for evolutionary constraint identification

Introducing phylogenetic tree-based training for nucleotide evolution modeling

Enhancing prediction of functionally disruptive variants using single sequences

Innovation

Methods, ideas, or system contributions that make the work stand out.

Phylogenetic tree modeling for genomic language

Multispecies alignment integrated in loss function

Single sequence prediction with transfer learning

🔎 Similar Papers

FGBERT: Function-Driven Pre-trained Gene Language Model for Metagenomics