Large language models reorganize representational geometry during in-context learning

📅 2026-05-16
🏛️ arXiv.org
📈 Citations: 1
Influential: 0
📄 PDF
🤖 AI Summary
This work investigates why large language models succeed in in-context learning on certain linearly separable binary classification tasks but fail on others. By constructing controlled binary classification tasks and integrating representational geometry analysis, causal interventions, and behavioral modeling, the study systematically examines the dynamic reorganization of representational structure during in-context learning. The findings reveal that successful in-context learning is not merely due to amplification of fixed label directions present in pretraining representations, but is instead driven by an active geometric restructuring of representations, which substantially enhances linear separability along task-relevant directions. Crucially, the geometry of pretrained representations imposes fundamental constraints on this process. Building on these insights, the authors propose a prototypical learner model, which for the first time identifies representational geometry reorganization as the key determinant of in-context learnability.
📝 Abstract
Large language models (LLMs) exhibit remarkable flexibility: they can adapt to novel tasks from in-context examples without any parameter updates, a capability known as in-context learning (ICL). Prior work on synthetic tasks has shown that ICL can implement specific algorithms, demonstrating architectural competence, and mechanistic analyses have identified key circuits that support this behavior. However, because in-context computation -- regardless of its algorithmic form -- relies on transformations in high-dimensional representation space, it remains unclear how the geometry of that space shapes ICL effectiveness. Motivated by the neuroscience view of classification as the untangling of neural representations, we hypothesize that ICL depends on the successful online untangling of task-relevant representations. To test this idea, we study how LLMs classify in-context examples whose labels are defined by the model's own internal representations with known structure. We show that ICL performance correlates systematically with the representational structure of the underlying classification task and that successful ICL is accompanied by geometric reorganization that increases online separability. We further find that LLM behavior is well described by a prototype-like algorithm that integrates evidence while reshaping representations to support classification. These findings offer a geometric account of ICL in pretrained LLMs, establish representational geometry as a mechanistic constraint on ICL, and quantify the gap between what pretrained representations afford and what in-context learning can exploit.
Problem

Research questions and friction points this paper is trying to address.

in-context learning
representational geometry
large language models
task learnability
linear separability
Innovation

Methods, ideas, or system contributions that make the work stand out.

in-context learning
representational geometry
large language models
geometric reorganization
prototype learning
🔎 Similar Papers