🤖 AI Summary
This study investigates the emergence mechanisms of interpretable semantic features in large language models (LLMs), characterizing their dynamic evolution across training time, Transformer layer depth, and model scale. Using sparse autoencoders (SAEs) for feature extraction, we integrate mechanistic interpretability techniques with cross-checkpoint, cross-layer, and cross-scale feature tracking to systematically identify critical emergence points for multi-domain semantic concepts. Our analysis reveals—for the first time—sharp temporal and scaling thresholds governing semantic feature emergence; further, it uncovers “unexpected reactivation” of early-layer features in deeper layers, challenging the conventional assumption of monotonic representational refinement. We construct the first three-dimensional emergence atlas—spanning time, spatial (layer-wise), and scale (parameter-count) dimensions—providing empirical grounding and a novel analytical framework for understanding how knowledge is formed and structured within LLMs.
📝 Abstract
This paper studies the emergence of interpretable categorical features within large language models (LLMs), analyzing their behavior across training checkpoints (time), transformer layers (space), and varying model sizes (scale). Using sparse autoencoders for mechanistic interpretability, we identify when and where specific semantic concepts emerge within neural activations. Results indicate clear temporal and scale-specific thresholds for feature emergence across multiple domains. Notably, spatial analysis reveals unexpected semantic reactivation, with early-layer features re-emerging at later layers, challenging standard assumptions about representational dynamics in transformer models.