Institution profile

Computer Science and Artificial Intelligence Laboratory (CSAIL), Massachusetts Institute of Technology (MIT)

Academic institutionnorthamerica · us
Official website
Research library21linked papers
Opportunities0open roles
Selected work

Representative Papers

Ptolemy: A Semantic Map of Exploratory Data Analysis

Sep 14, 2026

该研究通过构建一个名为Ptolemy的语义地图来解决探索性数据分析过程中难以追踪已分析内容的问题,利用结构化描述生成嵌入式位置表示每个分析步骤,以提高全局定位和局部比较能力。

0 citationsRead paper

Social.Wiki: A Web Held in Common

Aug 19, 2026

本文提出Social.Wiki系统,利用AI工具使非编程用户也能协作编辑社交网站,旨在解决网站所有者与用户利益不一致的问题。

0 citationsRead paper

Training Language Models to Explain Their Own Computations

Nov 11, 2025

This work investigates whether language models (LMs) can leverage “privileged access” to their internal computations to generate accurate, generalizable natural language explanations. Method: We introduce *self-explanation*—a novel paradigm wherein high-quality explanatory annotations are automatically generated via interpretability techniques (e.g., feature attribution, causal mediation analysis), and a pretrained LM is fine-tuned on only tens of thousands of such examples to produce explanations of feature encoding, activation-level causal structure, and input influence. Contribution/Results: Experiments demonstrate that self-explaining LMs significantly outperform strong external explainer models and generalize robustly to unseen queries with minimal training. Crucially, this is the first systematic empirical validation that privileged access to internal states yields substantial explanatory value—enabling scalable, low-cost model interpretation without requiring architectural modification or expensive human annotation.

0 citationsRead paper
Recent publications

Latest Papers

Ptolemy: A Semantic Map of Exploratory Data Analysis

Sep 14, 2026

该研究通过构建一个名为Ptolemy的语义地图来解决探索性数据分析过程中难以追踪已分析内容的问题,利用结构化描述生成嵌入式位置表示每个分析步骤,以提高全局定位和局部比较能力。

0 citationsRead paper

Social.Wiki: A Web Held in Common

Aug 19, 2026

本文提出Social.Wiki系统,利用AI工具使非编程用户也能协作编辑社交网站,旨在解决网站所有者与用户利益不一致的问题。

0 citationsRead paper

Training Language Models to Explain Their Own Computations

Nov 11, 2025

This work investigates whether language models (LMs) can leverage “privileged access” to their internal computations to generate accurate, generalizable natural language explanations. Method: We introduce *self-explanation*—a novel paradigm wherein high-quality explanatory annotations are automatically generated via interpretability techniques (e.g., feature attribution, causal mediation analysis), and a pretrained LM is fine-tuned on only tens of thousands of such examples to produce explanations of feature encoding, activation-level causal structure, and input influence. Contribution/Results: Experiments demonstrate that self-explaining LMs significantly outperform strong external explainer models and generalize robustly to unseen queries with minimal training. Crucially, this is the first systematic empirical validation that privileged access to internal states yields substantial explanatory value—enabling scalable, low-cost model interpretation without requiring architectural modification or expensive human annotation.

0 citationsRead paper