Depth-Wise Emergence of Prediction-Centric Geometry in Large Language Models
This work investigates how decoder-only large language models transform contextual information into predictive outputs along the depth dimension. By integrating geometric representation analysis with mechanistic interventions, the study reveals—for the first time—that angular components of deep-layer representations encode similarity structures aligned with predictive distributions, while norm components carry non-predictive contextual information. This finding establishes a mechanism-geometric account of the context-to-prediction transformation process. Building upon disentangled representations and an intervenable modeling framework, the research identifies structured geometric properties underlying prediction formation and demonstrates causal, selective control over token-level predictions.