MergeDNA: Context-aware Genome Modeling with Dynamic Tokenization through Token Merging
Genomic sequence modeling faces two fundamental challenges: highly non-uniform information density and the absence of natural minimal lexical units, rendering conventional single-base or static DNA tokenization approaches ill-suited to genomic structural complexity. To address this, we propose a unified framework integrating dynamic tokenization with context-aware pretraining. We introduce differentiable Token Merging—the first such application in genomics—to enable adaptive base-level aggregation. We further design a latent-variable Transformer with hierarchical attention, jointly enforcing local window constraints and global contextual modeling to support selective token identification and reconstruction. Our framework jointly optimizes tokenization and representation learning via two end-to-end objectives: merged-token reconstruction and adaptive masked modeling. Evaluated across three major DNA benchmarks and multi-omics tasks, our method consistently outperforms state-of-the-art tokenization strategies and large-scale DNA foundation models under both fine-tuning and zero-shot settings.