Institution profile

Martian

Research institution
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Ablation-Reversible Heads Don't Transfer: A Stress Test for Mechanistic Role Claims in Transformers

Jun 06, 2026

Existing interpretability methods often erroneously attribute specific computational roles to attention heads without verifying their generalization across diverse prompts. This work proposes the KID role classification framework and a three-stage analysis pipeline that combines activation patching with same-answer control conditions to expose widespread pseudo-semantic specificity in conventional attribution approaches. By integrating capability-selective screening (CSS), singular value decomposition (SVD), and activation transduction under matched controls, we systematically evaluate attention head functionality across multiple 7–8B instruction-tuned models. Our findings demonstrate that the majority of attention heads previously identified by standard methods as performing specific roles fail to consistently transfer their purported computational functions across different prompts, thereby challenging the dominant attribution paradigm in mechanistic interpretability.

0 citationsRead paper

Position: Require Frontier AI Labs To Release Small "Analog" Models

Oct 15, 2025

AI safety regulation often stifles innovation and increases compliance costs, necessitating resolution of the inherent trade-off between safety and technological advancement. Method: We propose an “Open Simulation Model” regulatory mechanism—requiring leading AI laboratories to distill knowledge from their state-of-the-art foundation models and construct functionally equivalent, compact, and interpretable open-source models, which are then publicly released. Contribution: This work pioneers the use of model distillation as a regulatory infrastructure component, enabling low-cost, high-throughput safety verification, interpretability analysis, and algorithmic auditing on lightweight models. Empirical results demonstrate that safety techniques developed on distilled models transfer effectively to their larger counterparts, substantially reducing regulatory overhead, fostering community-driven governance, and enhancing the transparency and governability of large language models. (149 words)

0 citationsRead paper

TinySQL: A Progressive Text-to-SQL Dataset for Mechanistic Interpretability Research

Mar 17, 2025

Mechanism interpretability research has long been hindered by the gap between synthetic “toy tasks” and real-world complexity. This paper proposes text-to-SQL generation as an ideal benchmark task—offering both formal syntactic structure and realistic semantic and compositional challenges. To this end, we introduce TinySQL, a progressively scaled synthetic dataset covering foundational to advanced SQL operations, and establish a dedicated evaluation platform spanning 33M–1B parameter models. We are the first to systematically apply mechanism interpretability methods to text-to-SQL. Our methodology integrates edge attribution patching, sparse autoencoders, circuit identification, and multi-scale comparative evaluation. This enables the precise localization of minimal functional circuits underlying SQL generation and reveals how such circuits dynamically reconfigure across query types. Critically, our analysis exposes fundamental limitations and biases of existing interpretability techniques, validates their utility in diagnosing model failure modes, and demonstrates their capacity to inform targeted dataset refinement.

0 citationsRead paper
Recent publications

Latest Papers

Ablation-Reversible Heads Don't Transfer: A Stress Test for Mechanistic Role Claims in Transformers

Jun 06, 2026

Existing interpretability methods often erroneously attribute specific computational roles to attention heads without verifying their generalization across diverse prompts. This work proposes the KID role classification framework and a three-stage analysis pipeline that combines activation patching with same-answer control conditions to expose widespread pseudo-semantic specificity in conventional attribution approaches. By integrating capability-selective screening (CSS), singular value decomposition (SVD), and activation transduction under matched controls, we systematically evaluate attention head functionality across multiple 7–8B instruction-tuned models. Our findings demonstrate that the majority of attention heads previously identified by standard methods as performing specific roles fail to consistently transfer their purported computational functions across different prompts, thereby challenging the dominant attribution paradigm in mechanistic interpretability.

0 citationsRead paper

Position: Require Frontier AI Labs To Release Small "Analog" Models

Oct 15, 2025

AI safety regulation often stifles innovation and increases compliance costs, necessitating resolution of the inherent trade-off between safety and technological advancement. Method: We propose an “Open Simulation Model” regulatory mechanism—requiring leading AI laboratories to distill knowledge from their state-of-the-art foundation models and construct functionally equivalent, compact, and interpretable open-source models, which are then publicly released. Contribution: This work pioneers the use of model distillation as a regulatory infrastructure component, enabling low-cost, high-throughput safety verification, interpretability analysis, and algorithmic auditing on lightweight models. Empirical results demonstrate that safety techniques developed on distilled models transfer effectively to their larger counterparts, substantially reducing regulatory overhead, fostering community-driven governance, and enhancing the transparency and governability of large language models. (149 words)

0 citationsRead paper

TinySQL: A Progressive Text-to-SQL Dataset for Mechanistic Interpretability Research

Mar 17, 2025

Mechanism interpretability research has long been hindered by the gap between synthetic “toy tasks” and real-world complexity. This paper proposes text-to-SQL generation as an ideal benchmark task—offering both formal syntactic structure and realistic semantic and compositional challenges. To this end, we introduce TinySQL, a progressively scaled synthetic dataset covering foundational to advanced SQL operations, and establish a dedicated evaluation platform spanning 33M–1B parameter models. We are the first to systematically apply mechanism interpretability methods to text-to-SQL. Our methodology integrates edge attribution patching, sparse autoencoders, circuit identification, and multi-scale comparative evaluation. This enables the precise localization of minimal functional circuits underlying SQL generation and reveals how such circuits dynamically reconfigure across query types. Critically, our analysis exposes fundamental limitations and biases of existing interpretability techniques, validates their utility in diagnosing model failure modes, and demonstrates their capacity to inform targeted dataset refinement.

0 citationsRead paper