Interpreting Protein Language Model Embeddings via Orthogonal Projection for Protein Fitness Prediction

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究使用正交投影技术去除蛋白质语言模型嵌入中的已知生化特征,以解释其对蛋白质适应性预测的影响,并量化这些特征的贡献。
📝 Abstract
Recently, there has been a growing adoption of protein language models (PLMs) in biomedical science. Their embeddings provide a rich numerical representation of protein sequences which achieve state-of-the-art performance on several downstream tasks including protein fitness prediction. However, PLM embeddings are not directly interpretable and, thereby, it remains unclear what features they encode. To gain insight into which biochemical properties of the protein are driving the prediction, we leverage an orthogonal projection technique that removes linear effects of known tabular features from embeddings and extend it to high-order and interaction effects. In this way, we remove the effects of interpretable biochemical features from PLM embeddings. In an ablation study, we show that this leads to a decrease in performance for a downstream classifier trained only on the embeddings to predict protein fitness. In an additional evaluation, we find that these biochemical features explain a substantial part of the variance in the predictions of this classifier. Hence, we can show that PLM embeddings encode patterns correlated with biochemical properties and quantify their contribution to predicting protein fitness. This computationally efficient approach is not limited to the features or embeddings considered here and is readily transferable to problem settings beyond protein fitness prediction.
Problem

Research questions and friction points this paper is trying to address.

protein language models
embeddings
interpretability
biochemical properties
fitness prediction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Orthogonal Projection
Protein Language Models (PLMs)
Interpretable Biochemical Features
Protein Fitness Prediction
P
Paulo Yanez Sarmiento
Hasso Plattner Institute, University of Potsdam
P
Pia Francesca Rissom
Hasso Plattner Institute, University of Potsdam; Bioinformatics and Machine Learning, Broad Institute of MIT and Harvard
M
Manuel Pfeuffer
School of Business and Economics, Humboldt University of Berlin
M
Marco Simnacher
School of Business and Economics, Humboldt University of Berlin
J
Jordan F. Safer
Bioinformatics and Machine Learning, Broad Institute of MIT and Harvard
Sumaiya Iqbal
Sumaiya Iqbal
Bioinformatics and Machine Learning, Broad Institute of MIT and Harvard
H
Henrike O. Heyne
Hasso Plattner Institute, University of Potsdam
Nadja Klein
Nadja Klein
Karlsruhe Institute of Technology
Statistical & Machine LearningBayesian Statistics & ComputingRegressionDeep Learning &
B
Bernhard Y. Renard
Hasso Plattner Institute, University of Potsdam