SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过SAEScientist-Bench评估AI代理能否使用稀疏自动编码器工具进行自主机制发现,以解决模型理解和安全对齐问题。
📝 Abstract
While research on recursive self-improvement (RSI) has predominantly automated model training pipelines, reliable autonomous development demands a missing pillar: post-hoc monitoring and auditing to understand what models learn and ensure safe alignment. Mechanistic interpretability tools are essential to bridge this gap, among which Sparse Autoencoders (SAEs) serve as a cornerstone by isolating interpretable features for model inspection and steering. In this paper, we introduce SAEScientist-Bench to evaluate whether AI agents can act as scientists utilizing SAE tools for autonomous mechanistic discovery. Given a target concept, an agent designs contrastive probes and navigates a Gemma Scope dictionary of 131K+ features in Gemma-2-9B-IT to discover the optimal feature, evaluated against curated expert reference features anchored on Neuronpedia across activation rank, concept selectivity on contrastive texts, and causal steering. Across 10 agent configurations and 20 tasks, frontier agents demonstrate genuine discovery capabilities and lead different evaluation dimensions, but remain well behind the expert baseline, approaching expert levels on separating target concepts from contrastive controls while lagging substantially in causal generation steering. Further analysis reveals that although agents can design contrasts to rule out spurious candidates, they frequently misinterpret experimental measurements. These results establish experimental model understanding as a measurable capability for closed-loop autonomous AI R&D. Our code is available at https://github.com/Trae1ounG/SAEScientist.
Problem

Research questions and friction points this paper is trying to address.

Recursive Self-Improvement
Mechanistic Interpretability
Sparse Autoencoders
Post-hoc Monitoring
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse Autoencoders
Mechanistic Interpretability
Autonomous AI R&D
Gemma-2-9B-IT
Contrastive Probes
Yuqiao Tan
Yuqiao Tan
Institute of Automation, Chinese Academy of Sciences
LLMs ReasoningLLMs Interpretability
S
Shizhu He
The Key Laboratory of Cognitive Intelligence, Institute of Automation, CAS; School of Artificial Intelligence, University of Chinese Academy of Sciences
Jun Zhao
Jun Zhao
School of Marine Sciences, Sun Yat-sen University
ocean opticsremote sensingnumerical modeling
K
Kang Liu
The Key Laboratory of Cognitive Intelligence, Institute of Automation, CAS; School of Artificial Intelligence, University of Chinese Academy of Sciences