Next-token pretraining implies in-context learning

📅 2025-05-23
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This paper investigates how standard self-supervised next-token prediction pretraining inherently induces in-distribution in-context learning (ICL), reframing it as a necessary consequence rather than an emergent phenomenon. Method: Leveraging an information-theoretic framework, we rigorously prove that minimizing prediction loss on non-ergodic token sequences necessitates implicit modeling of contextual dependencies—thereby guaranteeing ICL capability—and establish a precise mathematical coupling between ICL performance and the structural properties of the pretraining task. Our approach integrates information-theoretic analysis, synthetic data experiments, dynamical modeling of induction heads, and empirical validation of loss phase transitions and power-law scaling. Contribution/Results: We reproduce the phase transition in induction head emergence and quantitatively predict and verify ICL dynamics across data distributions with varying correlation structures. These results formally establish ICL as an intrinsic, provable property of next-token prediction pretraining.

Technology Category

Application Category

📝 Abstract
We argue that in-context learning (ICL) predictably arises from standard self-supervised next-token pretraining, rather than being an exotic emergent property. This work establishes the foundational principles of this emergence by focusing on in-distribution ICL, demonstrating how models necessarily adapt to context when trained on token sequences, especially from non-ergodic sources. Our information-theoretic framework precisely predicts these in-distribution ICL dynamics (i.e., context-dependent loss reduction). We verify this with experiments using synthetic datasets of differing types of correlational structure, reproducing characteristic phenomena like phase transitions in training loss for induction head formation and power-law scaling of in-context loss. We further show that a model's in-context performance on any task is mathematically coupled to the ensemble of tasks seen in pretraining, offering a fundamental explanation, grounded in architecture- and modality-independent principles, for such inference-time learning.
Problem

Research questions and friction points this paper is trying to address.

Explains in-context learning from next-token pretraining
Demonstrates context adaptation in non-ergodic token sequences
Links in-context performance to pretraining task ensemble
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-supervised next-token pretraining enables ICL
Information-theoretic framework predicts ICL dynamics
Pretraining task ensemble determines inference performance
🔎 Similar Papers
No similar papers found.
P
P. Riechers
Simplex, Astera Institute
H
Henry R. Bigelow
Simplex, Astera Institute
E
Eric A. Alt
Simplex, Astera Institute
A
Adam Shai
Simplex, Astera Institute