Decoding Translation-Related Functional Sequences in 5'UTRs Using Interpretable Deep Learning Models

📅 2025-07-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing 5′UTR translation efficiency prediction models suffer from fixed-length input constraints and limited interpretability. To address these limitations, we propose UTR-STCNet—a novel, interpretable deep learning architecture designed for variable-length sequence modeling. It integrates saliency-aware token clustering with a lightweight saliency-guided Transformer to enable multi-scale semantic aggregation and capture both local and long-range dependencies. By eliminating the need for sequence truncation, UTR-STCNet balances computational efficiency with biological interpretability. On three benchmark datasets, UTR-STCNet consistently outperforms state-of-the-art methods in predicting ribosomal load. Moreover, it successfully identifies key regulatory motifs—including upstream AUGs and Kozak sequences—demonstrating its capacity for biologically meaningful pattern discovery. This work establishes a new paradigm for functional interpretation and rational design of 5′UTRs.

Technology Category

Application Category

📝 Abstract
Understanding how 5' untranslated regions (5'UTRs) regulate mRNA translation is critical for controlling protein expression and designing effective therapeutic mRNAs. While recent deep learning models have shown promise in predicting translational efficiency from 5'UTR sequences, most are constrained by fixed input lengths and limited interpretability. We introduce UTR-STCNet, a Transformer-based architecture for flexible and biologically grounded modeling of variable-length 5'UTRs. UTR-STCNet integrates a Saliency-Aware Token Clustering (SATC) module that iteratively aggregates nucleotide tokens into multi-scale, semantically meaningful units based on saliency scores. A Saliency-Guided Transformer (SGT) block then captures both local and distal regulatory dependencies using a lightweight attention mechanism. This combined architecture achieves efficient and interpretable modeling without input truncation or increased computational cost. Evaluated across three benchmark datasets, UTR-STCNet consistently outperforms state-of-the-art baselines in predicting mean ribosome load (MRL), a key proxy for translational efficiency. Moreover, the model recovers known functional elements such as upstream AUGs and Kozak motifs, highlighting its potential for mechanistic insight into translation regulation.
Problem

Research questions and friction points this paper is trying to address.

Understanding 5'UTR regulation of mRNA translation
Improving interpretability in deep learning models for 5'UTRs
Predicting translational efficiency without input length constraints
Innovation

Methods, ideas, or system contributions that make the work stand out.

Transformer-based model for variable-length 5'UTRs
Saliency-Aware Token Clustering for interpretability
Lightweight attention captures local and distal dependencies
🔎 Similar Papers
No similar papers found.
Y
Yuxi Lin
Guangzhou National Laboratory, Guangzhou, Guangdong, China
Y
Yaxue Fang
Guangzhou National Laboratory, Guangzhou, Guangdong, China
Z
Zehong Zhang
Guangzhou National Laboratory, Guangzhou, Guangdong, China
Z
Zhouwu Liu
Guangzhou National Laboratory, Guangzhou, Guangdong, China
S
Siyun Zhong
Guangzhou National Laboratory, Guangzhou, Guangdong, China
Fulong Yu
Fulong Yu
Guangzhou National Laboratory
computational biologygenomicsbioinformaticsdeep learning