A Theoretical Framework for Masked Pretraining (MPT)

📅 2026-09-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入新的理论框架分析了掩码预训练(MPT)的工作机制,提出了均匀性增强的MPT损失(U-MPT)以解决维度坍缩问题,并提出了一种新的掩码策略来提高下游任务性能。
📝 Abstract
Recently, Masked Pretraining (MPT) based on reconstruction pretraining tasks has risen to a promising self-supervised learning paradigm across various domains and achieves remarkable performance in multiple downstream tasks. However, the theoretical understanding of the working mechanism behind MPT is still limited. In this paper, we introduce a new theoretical framework to analyze MPT and understand the crucial role of masking in extracting meaningful representations. We establish theoretical connections between MPT and another popular self-supervised paradigm: contrastive learning. We prove that the masking technique implicitly creates positive pairs that are semantically similar and the reconstruction loss pulls them together in the feature space. Besides, as a result of the implicit alignment, we point out the dimensional collapse issue of MPT and propose a Uniformity-enhanced MPT (U-MPT) loss that can effectively address this issue and bring significant improvements in downstream tasks including linear evaluation, cross-dataset fine-tuning and out-of-distribution generalization on real-world data sets. Furthermore, we establish downstream guarantees of U-MPT and theoretically analyze the influence of masking strategies. Based on the theoretical analysis, we propose a new masking strategy which enhances the downstream performance of MPT and explains current improvements of masking strategies with our theoretical perspective.
Problem

Research questions and friction points this paper is trying to address.

Masked Pretraining
theoretical understanding
dimensional collapse
Innovation

Methods, ideas, or system contributions that make the work stand out.

Masked Pretraining
theoretical framework
dimensional collapse
Uniformity-enhanced MPT
masking strategy
💼 Related Jobs
No related jobs found.
Q
Qi Zhang
State Key Lab of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University
R
Runyu Zhou
State Key Lab of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University
Yifei Wang
Yifei Wang
MIT
Machine LearningSelf-supervised LearningLanguage ModelsSparsitySafety
Yisen Wang
Yisen Wang
Assistant Professor, Peking University
Machine LearningSelf-Supervised LearningLarge Language ModelsSafety