Learning Task-Specific Antibody Representations via Function-Aware Masking

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过功能感知遮蔽方法改进抗体语言模型预训练,针对特定任务优化遮蔽策略,提高结构和CDR相关任务的表现。
📝 Abstract
Antibody-specific language models pretrained via masked language modeling (MLM) learn representations that are critical for downstream sequence design and property prediction tasks. Yet, the corruption process itself is rarely leveraged as a source of inductive bias during pretraining. While preferentially masking complementarity-determining regions (CDRs) improves binding-related predictions, antibodies possess diverse biological priors over a variety of functions. Herein, we introduce function-aware masking, a family of pretraining algorithms that align mask placement with specific functional priors (e.g., from IMGT annotations or structure predictions) to shape the learned representation space. We show that these specialist masking strategies significantly improve performance on their respective objectives, yielding up to a 14% gain on structure-related tasks and up to a 5.9x improvement on CDR-related tasks. To further improve performance across multiple functional axes, we develop hybrid masking strategies that integrate multiple priors, balancing reconstruction over binding, structural, and biophysical objectives. Our results demonstrate that informed mask placement provides a parameter-free mechanism for imposing functional inductive biases in antibody language model training.
Problem

Research questions and friction points this paper is trying to address.

antibody language models
masked language modeling
functional inductive biases
complementarity-determining regions
pretraining
Innovation

Methods, ideas, or system contributions that make the work stand out.

function-aware masking
specialist masking strategies
hybrid masking strategies
inductive bias