Institution profile

Simplex

Industry researchnorthamerica · us
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Rank-1 LoRAs Encode Interpretable Reasoning Signals

Nov 10, 2025

The internal mechanisms underlying performance improvements in reasoning models remain poorly understood. Method: We perform lightweight adaptation of Qwen-2.5-32B-Instruct using rank-1 LoRA, coupled with sparse autoencoder-based analysis of model activations, to uncover interpretable reasoning features embedded in low-rank adapters. Contribution: We demonstrate that minimal parameter perturbations—specifically, single-rank updates—are sufficient to elicit fine-grained, semantically homogeneous reasoning capabilities; their activation patterns are as interpretable as those of individual MLP neurons. On mainstream reasoning benchmarks (e.g., GSM8K, MMLU, HumanEval), our method recovers 73–90% of full fine-tuning performance. This work provides the first empirical evidence that complex reasoning abilities can be efficiently triggered by low-dimensional parameter changes, establishing a new paradigm for mechanistic interpretability that is both computationally efficient and highly interpretable.

0 citationsRead paper

Next-token pretraining implies in-context learning

May 23, 2025

This paper investigates how standard self-supervised next-token prediction pretraining inherently induces in-distribution in-context learning (ICL), reframing it as a necessary consequence rather than an emergent phenomenon. Method: Leveraging an information-theoretic framework, we rigorously prove that minimizing prediction loss on non-ergodic token sequences necessitates implicit modeling of contextual dependencies—thereby guaranteeing ICL capability—and establish a precise mathematical coupling between ICL performance and the structural properties of the pretraining task. Our approach integrates information-theoretic analysis, synthetic data experiments, dynamical modeling of induction heads, and empirical validation of loss phase transitions and power-law scaling. Contribution/Results: We reproduce the phase transition in induction head emergence and quantitatively predict and verify ICL dynamics across data distributions with varying correlation structures. These results formally establish ICL as an intrinsic, provable property of next-token prediction pretraining.

0 citationsRead paper

Constrained belief updates explain geometric structures in transformer representations

Feb 04, 2025

This work investigates the emergent computational structures in Transformers performing next-token prediction and their explanatory mechanisms for representational geometric features. Method: We propose a theoretical framework of “architecture-constrained parallel Bayesian belief updating,” unifying optimal prediction principles with mechanistic interpretability. Leveraging hidden Markov model (HMM) construction, probability simplex analysis, attention inverse modeling, and constraint-based refinement of optimal prediction equations, we quantitatively predict attention distributions, OV-circuit vector orientations, and embedding manifold geometry. Contribution/Results: Our framework rigorously derives the geometric structure of attention patterns, OV-circuit vectors, and token embeddings, establishing their formal correspondence to Bayesian inference. On controlled HMM tasks, it successfully reproduces and explains canonical geometric representations—including cyclic dynamics and low-dimensional manifolds—demonstrating both quantitative accuracy and mechanistic interpretability of the theoretical predictions.

0 citationsRead paper
Recent publications

Latest Papers

Rank-1 LoRAs Encode Interpretable Reasoning Signals

Nov 10, 2025

The internal mechanisms underlying performance improvements in reasoning models remain poorly understood. Method: We perform lightweight adaptation of Qwen-2.5-32B-Instruct using rank-1 LoRA, coupled with sparse autoencoder-based analysis of model activations, to uncover interpretable reasoning features embedded in low-rank adapters. Contribution: We demonstrate that minimal parameter perturbations—specifically, single-rank updates—are sufficient to elicit fine-grained, semantically homogeneous reasoning capabilities; their activation patterns are as interpretable as those of individual MLP neurons. On mainstream reasoning benchmarks (e.g., GSM8K, MMLU, HumanEval), our method recovers 73–90% of full fine-tuning performance. This work provides the first empirical evidence that complex reasoning abilities can be efficiently triggered by low-dimensional parameter changes, establishing a new paradigm for mechanistic interpretability that is both computationally efficient and highly interpretable.

0 citationsRead paper

Next-token pretraining implies in-context learning

May 23, 2025

This paper investigates how standard self-supervised next-token prediction pretraining inherently induces in-distribution in-context learning (ICL), reframing it as a necessary consequence rather than an emergent phenomenon. Method: Leveraging an information-theoretic framework, we rigorously prove that minimizing prediction loss on non-ergodic token sequences necessitates implicit modeling of contextual dependencies—thereby guaranteeing ICL capability—and establish a precise mathematical coupling between ICL performance and the structural properties of the pretraining task. Our approach integrates information-theoretic analysis, synthetic data experiments, dynamical modeling of induction heads, and empirical validation of loss phase transitions and power-law scaling. Contribution/Results: We reproduce the phase transition in induction head emergence and quantitatively predict and verify ICL dynamics across data distributions with varying correlation structures. These results formally establish ICL as an intrinsic, provable property of next-token prediction pretraining.

0 citationsRead paper

Constrained belief updates explain geometric structures in transformer representations

Feb 04, 2025

This work investigates the emergent computational structures in Transformers performing next-token prediction and their explanatory mechanisms for representational geometric features. Method: We propose a theoretical framework of “architecture-constrained parallel Bayesian belief updating,” unifying optimal prediction principles with mechanistic interpretability. Leveraging hidden Markov model (HMM) construction, probability simplex analysis, attention inverse modeling, and constraint-based refinement of optimal prediction equations, we quantitatively predict attention distributions, OV-circuit vector orientations, and embedding manifold geometry. Contribution/Results: Our framework rigorously derives the geometric structure of attention patterns, OV-circuit vectors, and token embeddings, establishing their formal correspondence to Bayesian inference. On controlled HMM tasks, it successfully reproduces and explains canonical geometric representations—including cyclic dynamics and low-dimensional manifolds—demonstrating both quantitative accuracy and mechanistic interpretability of the theoretical predictions.

0 citationsRead paper