Predictive Concept Decoders: Training Scalable End-to-End Interpretability Assistants
Interpretability of neural network internal activations has long been constrained by hand-crafted assumptions and scalability limitations of surrogate models. Method: We propose the first end-to-end trainable interpretability assistant that frames interpretability as a prediction task: a sparse concept encoder—acting as a communication bottleneck—maps internal activations to data-driven, natural-language concepts, while an autoregressive decoder directly predicts model behavior. Our approach employs a two-stage paradigm—self-supervised pretraining followed by instruction fine-tuning—and introduces an automatic evaluation metric (auto-interp score) to optimize bottleneck quality. Results: Experiments demonstrate significant improvements over baselines across diverse tasks—including jailbreak detection, implicit prompt identification, latent concept injection, and user attribute inference. The learned concept representations exhibit strong cross-task generalization, and both bottleneck quality and downstream performance scale consistently with data volume.