Institution profile

Flagship Pioneering

Industry researchnorthamerica · us
Official website
Research library5linked papers
Opportunities26open roles
Selected work

Representative Papers

PairSAE: Mechanistic Interpretability from Pair Representations in Protein Co-Folding

Jun 25, 2026

Existing sparse autoencoders struggle to effectively interpret pairwise representations in Pairformer-like protein co-folding models, often leading to feature explosion and an inability to model the joint distribution of sequence and pairwise features. This work proposes PairSAE, a novel framework that, for the first time, employs N-mode SVD to compress pairwise tensors into token-centric interaction roles and introduces a shared sparse autoencoder to jointly reconstruct both sequence and pairwise representations. By circumventing the quadratic growth inherent in conventional sparse autoencoders along pairwise dimensions, PairSAE extracts highly interpretable features on the PLINDER complex that align closely with UniProt functional annotations and accurately predicts Boltz-2 binding affinities, thereby uncovering structurally meaningful biological concepts learned internally by the model.

0 citationsRead paper

Dynamics of the Transformer Residual Stream: Coupling Spectral Geometry to Network Topology

May 13, 2026

The interlayer propagation mechanism of residual streams in large language models and their underlying spectral geometric structure remain poorly understood. This work treats Transformer depth as discrete time and models the residual stream as a dynamical system, integrating full Jacobian eigendecomposition, spectral geometry, graph community detection, and non-normal operator theory. It reveals, for the first time, a monotonic spectral gradient in trained models—transitioning from non-normal, rotation-dominated dynamics to near-symmetric behavior across layers. The study demonstrates that low-rank bottlenecks and the topological position of graph communities jointly govern perturbation amplification or suppression. These findings are validated across three production-scale large language models, showing that the observed spectral gradient and dimensional collapse are emergent properties of training, and that local operator types can predict perturbation propagation behavior.

0 citationsRead paper

Relaxed Sequence Sampling for Diverse Protein Design

Oct 27, 2025

Existing protein design methods (e.g., RSO) rely on single-path gradient optimization and neglect constraints inherent in sequence space, resulting in low sequence diversity and poor structural designability. To address this, we propose RSS, a Markov Chain Monte Carlo (MCMC)-based framework that jointly incorporates AlphaFold2’s structural prediction and ESM2’s evolutionary prior within a continuous logit-space energy function. RSS synergistically integrates gradient-guided sampling with language-model-informed “jumps” to simultaneously optimize structural accuracy and biological plausibility of sequences. Compared to RSO, RSS achieves a fivefold improvement in structural designability and a two- to threefold increase in sequence diversity at comparable computational cost. This substantially expands the tractable region of the protein design space while preserving physicochemical and evolutionary realism.

0 citationsRead paper

Flash Invariant Point Attention

May 16, 2025

In structural biology, invariant point attention (IPA) suffers from quadratic computational complexity, hindering its application to long protein/RNA sequences. This work introduces FlashIPA—the first algorithm reformulating IPA into a linear-complexity, hardware-efficient FlashAttention paradigm. By factorizing the attention mechanism, optimizing geometric tensor operations, and incorporating structure-specific embeddings, FlashIPA achieves strict linear scaling in both GPU memory consumption and runtime. Experiments demonstrate that FlashIPA enables end-to-end structure generation for sequences comprising thousands of residues, eliminating length constraints even when retraining generative models. Crucially, it matches or exceeds standard IPA in modeling accuracy while reducing GPU memory usage by 58% and accelerating inference by 2.3×. This work breaks IPA’s fundamental sequence-length bottleneck for the first time, establishing a scalable foundation for geometric modeling of large biomolecular systems.

0 citationsRead paper

Flow-of-Options: Diversified and Improved LLM Reasoning by Thinking Through Options

Feb 18, 2025

To address inherent biases in large language model (LLM) reasoning, this paper proposes Flow-of-Options (FoO), a framework that explicitly models diverse reasoning paths to enhance robustness and task adaptability. FoO introduces a novel “compressed interpretable representation” mechanism to enforce reasoning diversity and integrates case-based long-term memory for traceable solution generation and generalizable adaptation. The method unifies option-flow modeling, case-based reasoning, agent architecture, and multi-task adaptation to support autonomous machine learning (AutoML) task solving. Experiments demonstrate performance gains of 37.4%–69.2% on data science and therapeutic chemistry benchmarks, with per-task inference cost under $1. Furthermore, FoO successfully generalizes to reinforcement learning and image generation domains, validating its cross-modal applicability and scalability.

0 citationsRead paper
Recent publications

Latest Papers

PairSAE: Mechanistic Interpretability from Pair Representations in Protein Co-Folding

Jun 25, 2026

Existing sparse autoencoders struggle to effectively interpret pairwise representations in Pairformer-like protein co-folding models, often leading to feature explosion and an inability to model the joint distribution of sequence and pairwise features. This work proposes PairSAE, a novel framework that, for the first time, employs N-mode SVD to compress pairwise tensors into token-centric interaction roles and introduces a shared sparse autoencoder to jointly reconstruct both sequence and pairwise representations. By circumventing the quadratic growth inherent in conventional sparse autoencoders along pairwise dimensions, PairSAE extracts highly interpretable features on the PLINDER complex that align closely with UniProt functional annotations and accurately predicts Boltz-2 binding affinities, thereby uncovering structurally meaningful biological concepts learned internally by the model.

0 citationsRead paper

Dynamics of the Transformer Residual Stream: Coupling Spectral Geometry to Network Topology

May 13, 2026

The interlayer propagation mechanism of residual streams in large language models and their underlying spectral geometric structure remain poorly understood. This work treats Transformer depth as discrete time and models the residual stream as a dynamical system, integrating full Jacobian eigendecomposition, spectral geometry, graph community detection, and non-normal operator theory. It reveals, for the first time, a monotonic spectral gradient in trained models—transitioning from non-normal, rotation-dominated dynamics to near-symmetric behavior across layers. The study demonstrates that low-rank bottlenecks and the topological position of graph communities jointly govern perturbation amplification or suppression. These findings are validated across three production-scale large language models, showing that the observed spectral gradient and dimensional collapse are emergent properties of training, and that local operator types can predict perturbation propagation behavior.

0 citationsRead paper

Relaxed Sequence Sampling for Diverse Protein Design

Oct 27, 2025

Existing protein design methods (e.g., RSO) rely on single-path gradient optimization and neglect constraints inherent in sequence space, resulting in low sequence diversity and poor structural designability. To address this, we propose RSS, a Markov Chain Monte Carlo (MCMC)-based framework that jointly incorporates AlphaFold2’s structural prediction and ESM2’s evolutionary prior within a continuous logit-space energy function. RSS synergistically integrates gradient-guided sampling with language-model-informed “jumps” to simultaneously optimize structural accuracy and biological plausibility of sequences. Compared to RSO, RSS achieves a fivefold improvement in structural designability and a two- to threefold increase in sequence diversity at comparable computational cost. This substantially expands the tractable region of the protein design space while preserving physicochemical and evolutionary realism.

0 citationsRead paper

Flash Invariant Point Attention

May 16, 2025

In structural biology, invariant point attention (IPA) suffers from quadratic computational complexity, hindering its application to long protein/RNA sequences. This work introduces FlashIPA—the first algorithm reformulating IPA into a linear-complexity, hardware-efficient FlashAttention paradigm. By factorizing the attention mechanism, optimizing geometric tensor operations, and incorporating structure-specific embeddings, FlashIPA achieves strict linear scaling in both GPU memory consumption and runtime. Experiments demonstrate that FlashIPA enables end-to-end structure generation for sequences comprising thousands of residues, eliminating length constraints even when retraining generative models. Crucially, it matches or exceeds standard IPA in modeling accuracy while reducing GPU memory usage by 58% and accelerating inference by 2.3×. This work breaks IPA’s fundamental sequence-length bottleneck for the first time, establishing a scalable foundation for geometric modeling of large biomolecular systems.

0 citationsRead paper

Flow-of-Options: Diversified and Improved LLM Reasoning by Thinking Through Options

Feb 18, 2025

To address inherent biases in large language model (LLM) reasoning, this paper proposes Flow-of-Options (FoO), a framework that explicitly models diverse reasoning paths to enhance robustness and task adaptability. FoO introduces a novel “compressed interpretable representation” mechanism to enforce reasoning diversity and integrates case-based long-term memory for traceable solution generation and generalizable adaptation. The method unifies option-flow modeling, case-based reasoning, agent architecture, and multi-task adaptation to support autonomous machine learning (AutoML) task solving. Experiments demonstrate performance gains of 37.4%–69.2% on data science and therapeutic chemistry benchmarks, with per-task inference cost under $1. Furthermore, FoO successfully generalizes to reinforcement learning and image generation domains, validating its cross-modal applicability and scalability.

0 citationsRead paper