Large Language Models Develop Belief State Geometry In-Context

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过在隐藏马尔可夫模型数据上测试大语言模型,揭示了其上下文学习能力背后的信念状态几何表示,并证明了这些表示支持近似最优贝叶斯预测。
📝 Abstract
Large language models (LLMs) trained on next-token prediction exhibit remarkable in-context learning (ICL) abilities, yet the representations that support ICL remain poorly understood. We consider such representations in a controlled setting: prompting LLMs with data emitted from hidden Markov models (HMMs) and probing for the corresponding belief state -- the posterior distribution over the HMM's hidden states given the observed token history. Across six open-source LLMs prompted with data from 40 HMMs selected for non-trivial belief structure, we find that belief states are linearly decodable from residual stream activations, with peak probe $R^2$-values from 0.83-0.99 across HMM and LLM combinations, ranging from early to late layers. To establish functional relevance, we intervene directly on the probe-identified subspace via patching and steering, resulting in downstream prediction quality on the order of the untampered model, while controls degrade performance substantially. Together, these results provide representation-level evidence that ICL in open-source LLMs approximates optimal Bayesian prediction over a context-inferred generative model. More broadly, our findings extend prior results linking input-distribution structure to activation geometry: from toy networks trained explicitly on HMM data to production-scale LLMs.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
In-Context Learning
Hidden Markov Models
Belief State
Innovation

Methods, ideas, or system contributions that make the work stand out.

In-Context Learning
Belief State Decoding
Residual Stream Activations
Bayesian Prediction