Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过控制实验探讨了大型语言模型在预训练中如何通过辅助视角(知识的再表述)更有效地获取知识,揭示了数据多样性的重要性。
📝 Abstract
Gaps remain in our understanding of how large language models (LLMs) acquire knowledge during pre-training. We posit that auxiliary views, reformulations of knowledge, are causally helpful for learning. We design controlled experiments to isolate this. First, we confirm that repetition is necessary for acquisition and clarify that paraphrasing helps only at smaller batch sizes. Second, holding the token budget fixed, allocating tokens from document repetition to auxiliary views improves learning, counterintuitively, even for factual recall. Third, the effectiveness of auxiliary views is not contingent on the strength of the teacher model that generates them. Fourth, we identify forms of knowledge, contextual and foundational, that aid learning in the presence of prior knowledge gaps. Finally, we examine how these effects manifest mechanistically via layer-wise biases and compression. Together, our findings suggest that auxiliary representations of knowledge, which arise naturally in large pre-training corpora, are a key factor in the success of pre-training and offer a plausible explanation for why data diversity matters.
Problem

Research questions and friction points this paper is trying to address.

large language models
knowledge acquisition
pre-training
auxiliary views
Innovation

Methods, ideas, or system contributions that make the work stand out.

auxiliary views
knowledge acquisition
pre-training
large language models
data diversity