Evidence of Phase Transitions in Small Transformer-Based Language Models

📅 2025-11-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work investigates whether phase transitions—akin to those observed in large language models—emerge during the training of small Transformer language models, and whether they can be directly observed early in training on a linear time scale. To address this, the study employs character-level GPT-style models and introduces high-sensitivity dynamical probes: mean token length, fraction of correctly predicted tokens, and lexical diversity—monitored via Poisson and sub-Poisson statistical analyses to circumvent the insensitivity of conventional loss curves. Results demonstrate that sharp, unambiguous phase transitions occur robustly in the early training stage of small models, without requiring logarithmic time rescaling, and exhibit cross-scale universality. This constitutes the first empirical confirmation that phase transitions are an intrinsic, scale-invariant feature of language model training. The proposed probe metrics establish a novel analytical paradigm for studying training dynamics, enabling fine-grained, real-time characterization of emergent linguistic structure.

Technology Category

Application Category

📝 Abstract
Phase transitions have been proposed as the origin of emergent abilities in large language models (LLMs), where new capabilities appear abruptly once models surpass critical thresholds of scale. Prior work, such as that of Wei et al., demonstrated these phenomena under model and data scaling, with transitions revealed after applying a log scale to training compute. In this work, we ask three complementary questions: (1) Are phase transitions unique to large models, or can they also be observed in small transformer-based language models? (2) Can such transitions be detected directly in linear training space, rather than only after log rescaling? and (3) Can these transitions emerge at early stages of training? To investigate, we train a small GPT-style transformer on a character-level corpus and analyze the evolution of vocabulary usage throughout training. We track the average word length, the number of correct versus incorrect words, and shifts in vocabulary diversity. Building on these measures, we apply Poisson and sub-Poisson statistics to quantify how words connect and reorganize. This combined analysis reveals a distinct transition point during training. Notably, these transitions are not apparent in standard loss or validation curves, but become visible through our vocabulary- and statistics-based probes. Our findings suggest that phase-transition reorganizations are a general feature of language model training, observable even in modest models, detectable directly in linear training space, and occurring surprisingly early as coherence emerges. This perspective provides new insight into the nonlinear dynamics of language model training and underscores the importance of tailored metrics for uncovering phase transition behaviors
Problem

Research questions and friction points this paper is trying to address.

Detecting phase transitions in small transformer language models during training
Identifying transitions directly in linear training space without log rescaling
Observing vocabulary reorganization at early training stages using statistical probes
Innovation

Methods, ideas, or system contributions that make the work stand out.

Observing phase transitions in small transformer models
Using vocabulary statistics to detect transition points
Identifying transitions early in linear training space
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
N
Noah Hong
Lynbrook High School, San Jose, CA 95129 USA
T
Tao Hong
Keysight Technologies, Santa Rosa, CA 95403 USA