Limitations of Automated Simulatability: LLM Simulators Can Bypass Explanations

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了自动模拟性评估中LLM模拟器可能绕过解释的问题,通过定性复制和扩展ConSim的实验,揭示了两个主要限制,并提出改进评估方法的建议。
📝 Abstract
Simulatability is an evaluation protocol for explanations that quantifies their usefulness by how well they help a user predict a task model's outputs. Since human evaluation is costly, automated simulatability replaces human explainees with LLM simulators, as proposed in ConSim (Poch\'e et al., 2025) for large-scale experiments. We qualitatively replicate and extend ConSim's ranking of explanation methods across the tested datasets, explanation families, and simulator LLMs, and identify two limitations. First, when class names are meaningful, simulators can obtain high simulatability by solving the classification task directly, without relying on the explanations. Second, class anonymization can reward explanations for leaking the hidden label mapping, a limitation we expose with a new classes-as-concepts baseline. These results are consistent with a shortcut hypothesis: in the tested settings, simulator predictions mainly rely on task priors, while explanations produce small changes. We derive recommendations for more robust automated simulatability evaluations.
Problem

Research questions and friction points this paper is trying to address.

Automated Simulatability
LLM Simulators
Explanations Bypass
Class Anonymization
Shortcut Hypothesis
Innovation

Methods, ideas, or system contributions that make the work stand out.

automated simulatability
LLM simulators
classification task
class anonymization
shortcut hypothesis
🔎 Similar Papers
No similar papers found.