SeqMaestro: From nucleotide sequences to biological hypotheses through interpretable machine learning

📅 2026-09-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
SeqMaestro通过可解释的机器学习模型从核苷酸序列提出生物学假设,解决了传统方法灵活性不足和深度学习模型难以解释的问题。
📝 Abstract
Nucleotide sequence analysis is central to problems spanning regulatory genomics, evolutionary biology, and phenotype prediction. Classical bioinformatics methods extract interpretable sequence properties such as motifs and k-mer composition, but their flexibility is limited. In contrast, modern deep learning models can learn powerful predictive representations directly from raw sequences, yet their internal representations and decision mechanisms are difficult to inspect. Interpretable machine learning methods (e.g., sparse linear models and decision trees) provide human-understandable representations of predictive relationships but are not designed to operate directly on nucleotide sequences. Here, we introduce SeqMaestro, a machine learning framework that proposes biological hypotheses from nucleotide sequences using interpretable models. Our solution is centered around a two-layer interface that connects nucleotide sequences with the broader ecosystem of interpretable machine learning. SeqMaestro uses this interface to fit diverse combinations of interpretable models, feature representations, and extraction strategies, leveraging variability across transparent models to identify robust biological signals and richer predictive relationships than feature importance alone can provide. The system also supports data transformation and cleaning, model fitting, hyperparameter tuning, reliability analysis, and synthesis of results into a contextualized written report. By providing these capabilities through a no-code workflow, SeqMaestro is designed to make interpretable sequence analysis accessible to researchers without requiring extensive programming or machine learning expertise. SeqMaestro thereby provides an accessible route from nucleotide sequences to biological hypotheses.
Problem

Research questions and friction points this paper is trying to address.

nucleotide sequences
interpretable machine learning
biological hypotheses
Innovation

Methods, ideas, or system contributions that make the work stand out.

interpretable machine learning
nucleotide sequences
biological hypotheses
two-layer interface
no-code workflow
E
Evgeny S. Saveliev
Department of Applied Mathematics and Theoretical Physics, University of Cambridge, Cambridge, UK
Krzysztof Kacprzyk
Krzysztof Kacprzyk
PhD student, University of Cambridge
AI for ScienceInterpretabilityAI in Medicine
C
Charlotte Capitanchik
The Francis Crick Institute, London, UK; UK Dementia Research Institute at King’s College London, London, UK
N
Neelanjan Mukherjee
Department of Biochemistry and Molecular Genetics, University of Colorado Anschutz Medical Campus, Aurora, CO, USA
K
Kate Matlin
Department of Biochemistry and Molecular Genetics, University of Colorado Anschutz Medical Campus, Aurora, CO, USA
R
Ryan Sheridan
Department of Biochemistry and Molecular Genetics, University of Colorado Anschutz Medical Campus, Aurora, CO, USA
Srinivas Ramachandran
Srinivas Ramachandran
Department of Biochemistry and Molecular Genetics, University of Colorado Anschutz Medical Campus, Aurora, CO, USA
Jernej Ule
Jernej Ule
The Francis Crick Institute, London, UK; UK Dementia Research Institute at King’s College London, London, UK
D
David L. Bentley
Department of Biochemistry and Molecular Genetics, University of Colorado Anschutz Medical Campus, Aurora, CO, USA
Mihaela van der Schaar
Mihaela van der Schaar
University of Cambridge, The Alan Turing Institute
machine learningML for healthcarecompression and streamingmulti-user networkinggame-theory