🤖 AI Summary
研究通过使用基于Sentence BERT嵌入的简单签名方法,识别数据集中隐藏偏见的原因,解决了教师模型向学生模型传递未知偏见的问题。
📝 Abstract
Recent work has shown that a teacher model can transfer a bias to a student through a dataset from which every explicit reference to that bias has been filtered out, and that no data-level defense reliably removes or detects it even when knowing what bias to look for. Aiming to shed light on the hidden traces of these biases, we show that a simple signature based on Sentence BERT embeddings can identify the topic of such a bias with a Matthews correlation coefficient of 0.83 if the teacher model used by the attacker is known and 0.46 if it is not. Additionally, we observe that different teacher models appear to express the same bias through different vocabulary.