Large Language Models Develop Belief State Geometry In-Context
研究通过在隐藏马尔可夫模型数据上测试大语言模型,揭示了其上下文学习能力背后的信念状态几何表示,并证明了这些表示支持近似最优贝叶斯预测。
研究通过在隐藏马尔可夫模型数据上测试大语言模型,揭示了其上下文学习能力背后的信念状态几何表示,并证明了这些表示支持近似最优贝叶斯预测。
The internal mechanisms underlying performance improvements in reasoning models remain poorly understood. Method: We perform lightweight adaptation of Qwen-2.5-32B-Instruct using rank-1 LoRA, coupled with sparse autoencoder-based analysis of model activations, to uncover interpretable reasoning features embedded in low-rank adapters. Contribution: We demonstrate that minimal parameter perturbations—specifically, single-rank updates—are sufficient to elicit fine-grained, semantically homogeneous reasoning capabilities; their activation patterns are as interpretable as those of individual MLP neurons. On mainstream reasoning benchmarks (e.g., GSM8K, MMLU, HumanEval), our method recovers 73–90% of full fine-tuning performance. This work provides the first empirical evidence that complex reasoning abilities can be efficiently triggered by low-dimensional parameter changes, establishing a new paradigm for mechanistic interpretability that is both computationally efficient and highly interpretable.
This paper investigates how standard self-supervised next-token prediction pretraining inherently induces in-distribution in-context learning (ICL), reframing it as a necessary consequence rather than an emergent phenomenon. Method: Leveraging an information-theoretic framework, we rigorously prove that minimizing prediction loss on non-ergodic token sequences necessitates implicit modeling of contextual dependencies—thereby guaranteeing ICL capability—and establish a precise mathematical coupling between ICL performance and the structural properties of the pretraining task. Our approach integrates information-theoretic analysis, synthetic data experiments, dynamical modeling of induction heads, and empirical validation of loss phase transitions and power-law scaling. Contribution/Results: We reproduce the phase transition in induction head emergence and quantitatively predict and verify ICL dynamics across data distributions with varying correlation structures. These results formally establish ICL as an intrinsic, provable property of next-token prediction pretraining.
This work investigates the emergent computational structures in Transformers performing next-token prediction and their explanatory mechanisms for representational geometric features. Method: We propose a theoretical framework of “architecture-constrained parallel Bayesian belief updating,” unifying optimal prediction principles with mechanistic interpretability. Leveraging hidden Markov model (HMM) construction, probability simplex analysis, attention inverse modeling, and constraint-based refinement of optimal prediction equations, we quantitatively predict attention distributions, OV-circuit vector orientations, and embedding manifold geometry. Contribution/Results: Our framework rigorously derives the geometric structure of attention patterns, OV-circuit vectors, and token embeddings, establishing their formal correspondence to Bayesian inference. On controlled HMM tasks, it successfully reproduces and explains canonical geometric representations—including cyclic dynamics and low-dimensional manifolds—demonstrating both quantitative accuracy and mechanistic interpretability of the theoretical predictions.
研究通过在隐藏马尔可夫模型数据上测试大语言模型,揭示了其上下文学习能力背后的信念状态几何表示,并证明了这些表示支持近似最优贝叶斯预测。
The internal mechanisms underlying performance improvements in reasoning models remain poorly understood. Method: We perform lightweight adaptation of Qwen-2.5-32B-Instruct using rank-1 LoRA, coupled with sparse autoencoder-based analysis of model activations, to uncover interpretable reasoning features embedded in low-rank adapters. Contribution: We demonstrate that minimal parameter perturbations—specifically, single-rank updates—are sufficient to elicit fine-grained, semantically homogeneous reasoning capabilities; their activation patterns are as interpretable as those of individual MLP neurons. On mainstream reasoning benchmarks (e.g., GSM8K, MMLU, HumanEval), our method recovers 73–90% of full fine-tuning performance. This work provides the first empirical evidence that complex reasoning abilities can be efficiently triggered by low-dimensional parameter changes, establishing a new paradigm for mechanistic interpretability that is both computationally efficient and highly interpretable.
This paper investigates how standard self-supervised next-token prediction pretraining inherently induces in-distribution in-context learning (ICL), reframing it as a necessary consequence rather than an emergent phenomenon. Method: Leveraging an information-theoretic framework, we rigorously prove that minimizing prediction loss on non-ergodic token sequences necessitates implicit modeling of contextual dependencies—thereby guaranteeing ICL capability—and establish a precise mathematical coupling between ICL performance and the structural properties of the pretraining task. Our approach integrates information-theoretic analysis, synthetic data experiments, dynamical modeling of induction heads, and empirical validation of loss phase transitions and power-law scaling. Contribution/Results: We reproduce the phase transition in induction head emergence and quantitatively predict and verify ICL dynamics across data distributions with varying correlation structures. These results formally establish ICL as an intrinsic, provable property of next-token prediction pretraining.
This work investigates the emergent computational structures in Transformers performing next-token prediction and their explanatory mechanisms for representational geometric features. Method: We propose a theoretical framework of “architecture-constrained parallel Bayesian belief updating,” unifying optimal prediction principles with mechanistic interpretability. Leveraging hidden Markov model (HMM) construction, probability simplex analysis, attention inverse modeling, and constraint-based refinement of optimal prediction equations, we quantitatively predict attention distributions, OV-circuit vector orientations, and embedding manifold geometry. Contribution/Results: Our framework rigorously derives the geometric structure of attention patterns, OV-circuit vectors, and token embeddings, establishing their formal correspondence to Bayesian inference. On controlled HMM tasks, it successfully reproduces and explains canonical geometric representations—including cyclic dynamics and low-dimensional manifolds—demonstrating both quantitative accuracy and mechanistic interpretability of the theoretical predictions.