MUtE: A Dual Framework for Concept Erasure and Counterfactual Interventions
该研究通过设计一种新的概念擦除函数,解决了模型中概念信息的偏移问题,并提出了一种框架以实现概念擦除和反事实生成。
该研究通过设计一种新的概念擦除函数,解决了模型中概念信息的偏移问题,并提出了一种框架以实现概念擦除和反事实生成。
为解决LLM校准评估问题,通过操纵logit_bias参数,提出一种新的真校准误差估计器,实现对黑盒基础模型的有效审计。
Traditional probabilistic models are typically trained using task-agnostic log-loss, which often yields propensity scores with large errors, high bias, and high variance in boundary regions—particularly detrimental in causal inference tasks such as inverse probability weighting. This work proposes a general framework that, for the first time, integrates the error structure of downstream tasks into the design of strictly proper scoring rules. By aligning the local curvature of the scoring rule with that of the target loss, the authors derive a closed-form loss function tailored for average treatment effect estimation, along with its associated canonical probability mapping, enabling end-to-end task-oriented training. The approach is compatible with both neural networks and gradient boosting models and consistently outperforms standard log-likelihood and covariate balancing methods across multiple causal inference benchmarks, substantially improving estimation accuracy and stability.
Existing collaborative filtering methods suffer from limited capability in fine-grained preference modeling and explanation generation, failing to jointly optimize rating prediction accuracy and personalized interpretability. To address this, we propose a multi-task explainable recommendation framework that jointly learns global and aspect-level user–item representations. A personalized attention mechanism is introduced to dynamically weight the importance of different aspects according to individual user preferences. Leveraging T5-small, the model simultaneously optimizes three objectives: overall rating prediction, aspect-level rating prediction, and personalized review generation. Extensive experiments on TripAdvisor and RateBeer demonstrate that our approach significantly outperforms strong baselines—particularly in generating semantically coherent, user-specific, high-quality natural language explanations. To the best of our knowledge, this is the first work to achieve deep joint modeling of precise rating prediction and fine-grained, natural-language-based explanations.
This paper addresses the problem of irreversibly erasing sensitive demographic attributes (e.g., gender, race) from neural representations while preserving semantic meaning to improve fairness in downstream NLP tasks. To this end, we propose LEOPARD: a density-matching-based orthogonal projection learning framework that achieves efficient, structure-preserving erasure of discrete concepts in nonlinear embedding spaces. LEOPARD explicitly aligns class-conditional feature distributions and controls projection rank to ensure both geometric fidelity and erasure rigor. Its key contribution lies in formulating concept erasure as a distribution alignment problem with local geometric constraints, thereby balancing strict attribute removal and high semantic retention. Evaluated on multiple NLP benchmarks, LEOPARD significantly outperforms state-of-the-art methods—achieving superior bias mitigation under deep nonlinear classifiers while maintaining competitive task performance.
该研究通过设计一种新的概念擦除函数,解决了模型中概念信息的偏移问题,并提出了一种框架以实现概念擦除和反事实生成。
为解决LLM校准评估问题,通过操纵logit_bias参数,提出一种新的真校准误差估计器,实现对黑盒基础模型的有效审计。
Traditional probabilistic models are typically trained using task-agnostic log-loss, which often yields propensity scores with large errors, high bias, and high variance in boundary regions—particularly detrimental in causal inference tasks such as inverse probability weighting. This work proposes a general framework that, for the first time, integrates the error structure of downstream tasks into the design of strictly proper scoring rules. By aligning the local curvature of the scoring rule with that of the target loss, the authors derive a closed-form loss function tailored for average treatment effect estimation, along with its associated canonical probability mapping, enabling end-to-end task-oriented training. The approach is compatible with both neural networks and gradient boosting models and consistently outperforms standard log-likelihood and covariate balancing methods across multiple causal inference benchmarks, substantially improving estimation accuracy and stability.
Existing collaborative filtering methods suffer from limited capability in fine-grained preference modeling and explanation generation, failing to jointly optimize rating prediction accuracy and personalized interpretability. To address this, we propose a multi-task explainable recommendation framework that jointly learns global and aspect-level user–item representations. A personalized attention mechanism is introduced to dynamically weight the importance of different aspects according to individual user preferences. Leveraging T5-small, the model simultaneously optimizes three objectives: overall rating prediction, aspect-level rating prediction, and personalized review generation. Extensive experiments on TripAdvisor and RateBeer demonstrate that our approach significantly outperforms strong baselines—particularly in generating semantically coherent, user-specific, high-quality natural language explanations. To the best of our knowledge, this is the first work to achieve deep joint modeling of precise rating prediction and fine-grained, natural-language-based explanations.
This paper addresses the problem of irreversibly erasing sensitive demographic attributes (e.g., gender, race) from neural representations while preserving semantic meaning to improve fairness in downstream NLP tasks. To this end, we propose LEOPARD: a density-matching-based orthogonal projection learning framework that achieves efficient, structure-preserving erasure of discrete concepts in nonlinear embedding spaces. LEOPARD explicitly aligns class-conditional feature distributions and controls projection rank to ensure both geometric fidelity and erasure rigor. Its key contribution lies in formulating concept erasure as a distribution alignment problem with local geometric constraints, thereby balancing strict attribute removal and high semantic retention. Evaluated on multiple NLP benchmarks, LEOPARD significantly outperforms state-of-the-art methods—achieving superior bias mitigation under deep nonlinear classifiers while maintaining competitive task performance.