Tracing Stereotypes from Representation to Output in Multilingual LLMs

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究使用多种方法探究多语言大模型中刻板印象的内部机制及其对输出的影响,发现不同语言间存在差异且关键特征具有语言特异性。
📝 Abstract
Multilingual LLMs show stereotype-related behavior that varies across languages, but behavioral scores do not show where the relevant information is represented or how it affects the output. To investigate these internal mechanisms, we compare linear probing, attribution patching, sparse autoencoders (SAEs) and feature ablation in Llama-3.1-8B, Qwen3-8B, and Gemma-2-9B. Probe performance peaks substantially earlier than attribution in all three models, with a separation of 36-53% of model depth. Retained Llama-Scope features often match the social category on which they were selected and form recurring semantic families, but their lexical alignment and ablation effects vary across SAE suites. Only 6-18% of evaluated residual-stream features have language-agnostic effects under our criterion, and none are category-agnostic. Language-agnostic features have larger mean ablation effects in Llama-Scope, but this pattern does not repeat in the other SAE suites. Decodability, output influence, and cross-lingual ablation effects therefore need to be measured separately.
Problem

Research questions and friction points this paper is trying to address.

stereotypes
multilingual LLMs
representation
output
Innovation

Methods, ideas, or system contributions that make the work stand out.

linear probing
attribution patching
sparse autoencoders
feature ablation
language-agnostic features
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Ariun-Erdene Tumurchuluun
Saarland University, Saarland Informatics Campus
Y
Yusser Al Ghussin
Saarland University, Saarland Informatics Campus; German Research Center for Artificial Intelligence (DFKI)
Pinzhen Chen
Pinzhen Chen
University of Edinburgh
large language modelsLLM post-trainingmachine translationmultilinguality
Josef van Genabith
Josef van Genabith
DFKI German Research Center for Artificial Intelligence, Saarland University
Natural Language ProcessingMachine TranslationComputational LinguisticsComputational Semantics
K
Koel Dutta Chowdhury
University of Technology Nuremberg