Institution profile

Leap Laboratories

Academic institution
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Exemplar Partitioning for Mechanistic Interpretability

May 14, 2026

This work proposes an unsupervised Example Partitioning (EP) method to efficiently construct interpretable and computationally lightweight feature dictionaries for analyzing the internal mechanisms of large language models. EP leverages real samples from streaming activation data as Voronoi region anchors, eliminating the need to predefine dictionary size, and uniquely employs observed activations directly as both intervention directions and region representatives. This enables feature alignment across layers, models, and training stages while inherently supporting out-of-distribution detection. By integrating distance-threshold leader clustering with causal interventions, EP achieves superior performance on Gemma-2-2B, surpassing GemmaScope SAE’s AxBench AUROC (0.881) at only one-thousandth of the computational cost, retaining 97% probe accuracy, and exhibiting high consistency with SAE features in 20% of its regions.

0 citationsRead paper

Benchmarking the Discovery Engine

Jul 01, 2025

Current machine learning models suffer from limited interpretability and insufficient capacity to generate scientific insights. To address this, we propose Discovery Engine—a fully automated, end-to-end scientific discovery system that integrates multi-source heterogeneous data modeling with state-of-the-art interpretability techniques, including causal inference, concept activation mapping, and symbolic induction. We evaluate the framework across four domains—medicine, materials science, social sciences, and environmental science—using five published benchmark studies. Discovery Engine matches or substantially outperforms SOTA methods in predictive accuracy. Crucially, it systematically generates high-level scientific outputs: mechanistic explanations, empirically testable hypotheses, and actionable intervention strategies. Results demonstrate that the framework not only enhances model trustworthiness but also enables the practical realization of an “interpretability-driven scientific discovery” paradigm, establishing a new benchmark for automated scientific discovery.

0 citationsRead paper
Recent publications

Latest Papers

Exemplar Partitioning for Mechanistic Interpretability

May 14, 2026

This work proposes an unsupervised Example Partitioning (EP) method to efficiently construct interpretable and computationally lightweight feature dictionaries for analyzing the internal mechanisms of large language models. EP leverages real samples from streaming activation data as Voronoi region anchors, eliminating the need to predefine dictionary size, and uniquely employs observed activations directly as both intervention directions and region representatives. This enables feature alignment across layers, models, and training stages while inherently supporting out-of-distribution detection. By integrating distance-threshold leader clustering with causal interventions, EP achieves superior performance on Gemma-2-2B, surpassing GemmaScope SAE’s AxBench AUROC (0.881) at only one-thousandth of the computational cost, retaining 97% probe accuracy, and exhibiting high consistency with SAE features in 20% of its regions.

0 citationsRead paper

Benchmarking the Discovery Engine

Jul 01, 2025

Current machine learning models suffer from limited interpretability and insufficient capacity to generate scientific insights. To address this, we propose Discovery Engine—a fully automated, end-to-end scientific discovery system that integrates multi-source heterogeneous data modeling with state-of-the-art interpretability techniques, including causal inference, concept activation mapping, and symbolic induction. We evaluate the framework across four domains—medicine, materials science, social sciences, and environmental science—using five published benchmark studies. Discovery Engine matches or substantially outperforms SOTA methods in predictive accuracy. Crucially, it systematically generates high-level scientific outputs: mechanistic explanations, empirically testable hypotheses, and actionable intervention strategies. Results demonstrate that the framework not only enhances model trustworthiness but also enables the practical realization of an “interpretability-driven scientific discovery” paradigm, establishing a new benchmark for automated scientific discovery.

0 citationsRead paper