MUtE: A Dual Framework for Concept Erasure and Counterfactual Interventions

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究通过设计一种新的概念擦除函数,解决了模型中概念信息的偏移问题,并提出了一种框架以实现概念擦除和反事实生成。
📝 Abstract
Erasing concept-specific information from representations has been proven useful for mitigating bias or interpreting model decisions. The joint objective is to transform the original representations such that the target concept becomes unpredictable, while maximally preserving concept-unrelated information. In this work, we revisit the optimal bounds of concept erasure to derive a novel class of erasure functions that naturally induce a deterministic, dual counterfactual mapping. Bridging the gap between theoretical optimality and practical representation learning, we design an implementation that imposes a translational bias on counterfactual trajectories - a constraint that aligns with how many concepts geometrically manifest in modern language models. Our framework enables seamless navigation between concept erasure and counterfactual generation. We empirically demonstrate its efficacy in improving downstream algorithmic fairness and generating counterfactual texts.
Problem

Research questions and friction points this paper is trying to address.

concept erasure
counterfactual interventions
representation learning
algorithmic fairness
Innovation

Methods, ideas, or system contributions that make the work stand out.

concept erasure
counterfactual interventions
dual mapping
translational bias
representation learning
🔎 Similar Papers
A
Antoine Saillenfest
onepoint, 29 rue des Sablons, 75116 Paris (France)