Phantom transitions in language model fine-tuning

📅 2026-05-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the "silent failure" phenomenon in language model fine-tuning—where correct tokens fail to outcompete semantically similar alternatives despite a steadily decreasing cross-entropy loss—by introducing a novel analytical framework based on density matrices. The authors construct an order parameter that integrates predictive distributions with geometric overlap in token embedding space, decomposing prediction dynamics into signal and background drag components. This approach reveals, for the first time in non-orthogonal embedding spaces, two distinct failure mechanisms: kinematic and structural failures, and clarifies the nature of pseudo-phase transitions. Combining geometric embedding analysis, LoRA comparisons, and gradient-step-level dynamics monitoring, the study identifies a universal dimensionless quantity under full-parameter fine-tuning that accurately predicts the critical learning rate for leave-one-out architectures, achieving a prediction error of only 2.1%.
📝 Abstract
Fine-tuning a language model on contexts whose correct completion has a near-synonym competitor often fails silently. The cross-entropy loss decreases monotonically while the correct token never overtakes the competitor in rank. We study this regime across five transformer architectures spanning two families and a fivefold parameter range, on ten hand-selected near-synonym contexts. We instrument these failures with an order parameter combining the predicted distribution and pairwise embedding overlaps. It decomposes additively into a signal, tracking the model's commitment to the correct token over its nearest competitor, and a background drag, set by how the embedding bulk leaks probability into the score. This isolates two failure modes. In kinematic failure the signal stays small. In structural failure the drag actively worsens as fine-tuning proceeds. We observe sharp catapult-like jumps in the order parameter that resemble a phase transition. A central negative result organises the paper. The transitions are phantoms. The spontaneous-symmetry-breaking interpretation is ruled out by direct measurement. Catapult-like jumps still appear under LoRA fine-tuning with the token embedding matrix exactly unchanged during training, where no geometric phase transition is possible. The discontinuity lives entirely in the softmax readout. A small number of dimensionless quantities organise the trajectory across architectures. One is consistent across all five under full fine-tuning. A second sorts architectures into two classes by bulk embedding distribution and predicts LoRA sufficiency. As a blind test, the framework predicts the critical learning rate of a held-out architecture, not used to fit any parameter, to within 2.1% of a subsequent learning-rate sweep. Findings concern the near-synonym mechanism only and should not be extrapolated without recalibration.
Problem

Research questions and friction points this paper is trying to address.

fine-tuning failure
near-synonym competition
language models
phantom transitions
embedding geometry
Innovation

Methods, ideas, or system contributions that make the work stand out.

density-matrix
phantom transitions
geometric overlap
order parameter
LoRA fine-tuning
V
Vaibhav Prakash
Department of Physics, Mahindra University, Hyderabad, India
J
Jayasri Dontabhaktuni
Department of Physics, Mahindra University, Hyderabad, India