Institution profile

Transluce

Industry research
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Predictive Concept Decoders: Training Scalable End-to-End Interpretability Assistants

Dec 17, 2025

Interpretability of neural network internal activations has long been constrained by hand-crafted assumptions and scalability limitations of surrogate models. Method: We propose the first end-to-end trainable interpretability assistant that frames interpretability as a prediction task: a sparse concept encoder—acting as a communication bottleneck—maps internal activations to data-driven, natural-language concepts, while an autoregressive decoder directly predicts model behavior. Our approach employs a two-stage paradigm—self-supervised pretraining followed by instruction fine-tuning—and introduces an automatic evaluation metric (auto-interp score) to optimize bottleneck quality. Results: Experiments demonstrate significant improvements over baselines across diverse tasks—including jailbreak detection, implicit prompt identification, latent concept injection, and user attribute inference. The learned concept representations exhibit strong cross-task generalization, and both bottleneck quality and downstream performance scale consistently with data volume.

0 citationsRead paper

Training Language Models to Explain Their Own Computations

Nov 11, 2025

This work investigates whether language models (LMs) can leverage “privileged access” to their internal computations to generate accurate, generalizable natural language explanations. Method: We introduce *self-explanation*—a novel paradigm wherein high-quality explanatory annotations are automatically generated via interpretability techniques (e.g., feature attribution, causal mediation analysis), and a pretrained LM is fine-tuned on only tens of thousands of such examples to produce explanations of feature encoding, activation-level causal structure, and input influence. Contribution/Results: Experiments demonstrate that self-explaining LMs significantly outperform strong external explainer models and generalize robustly to unseen queries with minimal training. Crucially, this is the first systematic empirical validation that privileged access to internal states yields substantial explanatory value—enabling scalable, low-cost model interpretation without requiring architectural modification or expensive human annotation.

0 citationsRead paper
Recent publications

Latest Papers

Predictive Concept Decoders: Training Scalable End-to-End Interpretability Assistants

Dec 17, 2025

Interpretability of neural network internal activations has long been constrained by hand-crafted assumptions and scalability limitations of surrogate models. Method: We propose the first end-to-end trainable interpretability assistant that frames interpretability as a prediction task: a sparse concept encoder—acting as a communication bottleneck—maps internal activations to data-driven, natural-language concepts, while an autoregressive decoder directly predicts model behavior. Our approach employs a two-stage paradigm—self-supervised pretraining followed by instruction fine-tuning—and introduces an automatic evaluation metric (auto-interp score) to optimize bottleneck quality. Results: Experiments demonstrate significant improvements over baselines across diverse tasks—including jailbreak detection, implicit prompt identification, latent concept injection, and user attribute inference. The learned concept representations exhibit strong cross-task generalization, and both bottleneck quality and downstream performance scale consistently with data volume.

0 citationsRead paper

Training Language Models to Explain Their Own Computations

Nov 11, 2025

This work investigates whether language models (LMs) can leverage “privileged access” to their internal computations to generate accurate, generalizable natural language explanations. Method: We introduce *self-explanation*—a novel paradigm wherein high-quality explanatory annotations are automatically generated via interpretability techniques (e.g., feature attribution, causal mediation analysis), and a pretrained LM is fine-tuned on only tens of thousands of such examples to produce explanations of feature encoding, activation-level causal structure, and input influence. Contribution/Results: Experiments demonstrate that self-explaining LMs significantly outperform strong external explainer models and generalize robustly to unseen queries with minimal training. Crucially, this is the first systematic empirical validation that privileged access to internal states yields substantial explanatory value—enabling scalable, low-cost model interpretation without requiring architectural modification or expensive human annotation.

0 citationsRead paper