Institution profile

Intuition Machines, Inc.

Industry researchnorthamerica · us
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Where does Absolute Position come from in decoder-only Transformers?

Jun 04, 2026

This study investigates why decoder-only Transformers employing only relative positional encoding (RoPE) nonetheless exhibit absolute positional awareness. Through theoretical analysis and ablation experiments, the authors uncover that the key mechanism enabling absolute position leakage lies in the dynamic coupling between the softmax normalization term within the causal mask and the residual stream at position 0. They introduce the concept of an “attention sink” to stabilize token anchoring at this initial position. By integrating variants such as NTK scaling and sliding window attention, the work further examines how different components influence positional information propagation. Experiments demonstrate that replacing the BOS embedding reduces residual stream contributions in early queries by 40%, confirming that the attention sink conveys a deterministic fingerprint of the position-0 token, thereby explaining cross-input discrepancies in absolute positional behavior.

0 citationsRead paper

Where Pretraining writes and Alignment reads: the asymmetry of Transformer weight space

May 15, 2026

This work investigates the geometric imprints left by pretraining and alignment in the weight space of Transformer models and their underlying causes. Through subspace alignment analysis, gradient covariance modeling, optimization trajectory tracking, and rank-1 interventions, the study systematically uncovers an asymmetric update pattern between read and write paths: alignment updates concentrate along dominant directions in the read path, while the write path remains nearly isotropic. The authors propose an “anisotropic gradient accumulation” mechanism to explain this phenomenon and validate its efficacy via comparative objective control and causal intervention experiments. These findings offer a geometric perspective and theoretical foundation for understanding how alignment reshapes pretrained models.

0 citationsRead paper

The Phenomenology of Hallucinations

Mar 14, 2026

This work reveals that although language models internally encode uncertainty signals, these signals are weakly coupled to the output layer, preventing the model from abstaining and thereby generating hallucinations. For the first time, the study analyzes this limitation through the lens of representation geometry and topology, demonstrating that uncertainty manifests as fragmented structures in high-dimensional space, lacking a unified abstention attractor. The mechanism is validated across diverse model architectures using intrinsic dimension estimation, gradient and Fisher information probing, topological data analysis, and causal interventions. By directly injecting the internal uncertainty signal into the logits, the authors significantly restore the model’s ability to abstain, effectively suppressing hallucinations and offering a novel pathway toward reliable generation.

0 citationsRead paper
Recent publications

Latest Papers

Where does Absolute Position come from in decoder-only Transformers?

Jun 04, 2026

This study investigates why decoder-only Transformers employing only relative positional encoding (RoPE) nonetheless exhibit absolute positional awareness. Through theoretical analysis and ablation experiments, the authors uncover that the key mechanism enabling absolute position leakage lies in the dynamic coupling between the softmax normalization term within the causal mask and the residual stream at position 0. They introduce the concept of an “attention sink” to stabilize token anchoring at this initial position. By integrating variants such as NTK scaling and sliding window attention, the work further examines how different components influence positional information propagation. Experiments demonstrate that replacing the BOS embedding reduces residual stream contributions in early queries by 40%, confirming that the attention sink conveys a deterministic fingerprint of the position-0 token, thereby explaining cross-input discrepancies in absolute positional behavior.

0 citationsRead paper

Where Pretraining writes and Alignment reads: the asymmetry of Transformer weight space

May 15, 2026

This work investigates the geometric imprints left by pretraining and alignment in the weight space of Transformer models and their underlying causes. Through subspace alignment analysis, gradient covariance modeling, optimization trajectory tracking, and rank-1 interventions, the study systematically uncovers an asymmetric update pattern between read and write paths: alignment updates concentrate along dominant directions in the read path, while the write path remains nearly isotropic. The authors propose an “anisotropic gradient accumulation” mechanism to explain this phenomenon and validate its efficacy via comparative objective control and causal intervention experiments. These findings offer a geometric perspective and theoretical foundation for understanding how alignment reshapes pretrained models.

0 citationsRead paper

The Phenomenology of Hallucinations

Mar 14, 2026

This work reveals that although language models internally encode uncertainty signals, these signals are weakly coupled to the output layer, preventing the model from abstaining and thereby generating hallucinations. For the first time, the study analyzes this limitation through the lens of representation geometry and topology, demonstrating that uncertainty manifests as fragmented structures in high-dimensional space, lacking a unified abstention attractor. The mechanism is validated across diverse model architectures using intrinsic dimension estimation, gradient and Fisher information probing, topological data analysis, and causal interventions. By directly injecting the internal uncertainty signal into the logits, the authors significantly restore the model’s ability to abstain, effectively suppressing hallucinations and offering a novel pathway toward reliable generation.

0 citationsRead paper