Structural priors for data-efficient language learning

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过在非语言数据上预训练模型以减少对大量数据和计算资源的依赖,进而提高自然语言学习效率,但发现这对下游语言任务性能提升有限。
📝 Abstract
Efficient language learning requires methods to reduce the reliance on large data and computational resources. We investigate structural transfer: First training models on non-language data to induce useful priors for natural language. This approach is a form of weight initialization for multilingual language modeling. We evaluate transfer via next-token-prediction loss, weight shifts in the model, and downstream linguistic benchmarks. Several symbolic data types - notably music, probabilistic grammars, and cellular automata - yield lower language-modeling loss than random initialization. These gains coincide with smaller weight shifts during subsequent language training, suggesting that structural transfer positions models in a more favorable region of the parameter space. However, a lower loss does not translate consistently into better downstream linguistic performance, and transfer from non-language data is less efficient than additional language data. We conclude that non-language data can serve as a partial substitute for language data for the training objective of next-token prediction but does not reliably support broader linguistic generalization.
Problem

Research questions and friction points this paper is trying to address.

data-efficient
language learning
structural transfer
Innovation

Methods, ideas, or system contributions that make the work stand out.

structural transfer
weight initialization
next-token prediction
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Yana Veitsman
Yana Veitsman
Saarland University
multilingualitymechanistic interpretabilityrobust and efficient NLP
J
Jonas Mayer Martins
University of Göttingen, Germany
J
Jonathan Lautenschlager
University of Göttingen, Germany
Lisa Beinborn
Lisa Beinborn
Human-Centered Data Science, University of Göttingen