Distribution-aware Language Neuron Identification in Multilingual Large Language Models

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种新的方法,通过考虑激活分布的成对关系来更准确地识别多语言大模型中的语言特异性神经元。
📝 Abstract
Multilingual large language models (mLLMs) contain a small fraction of feed-forward neurons that are sensitive to particular languages, commonly termed language-specific neurons. Existing work measures language specificity using the entropy of each neuron's language-wise probabilities of being active, where a neuron is considered active when its activation value is positive. However, this approach may not fully capture the multilingual nature of mLLMs, where language representations are distributional and mutually related. We propose Distribution-aware Language Neuron selection, which leverages pairwise relationships between per-language activation distributions over the full activation range, including negative values. Specifically, we quantify each neuron's language specificity by clustering languages using pairwise overlap coefficients between their activation distributions. Across two mLLMs and two held-out corpora, our identifier more effectively isolates language-specific causal effects, yielding up to 4.9$\times$ higher on-target language damage per neuron while preserving off-target language performance.
Problem

Research questions and friction points this paper is trying to address.

language-specific neurons
multilingual large language models
activation distributions
language specificity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Distribution-aware Language Neuron selection
pairwise overlap coefficients
activation distributions
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
M
Minjun Kim
KAIST
I
Inho Won
KAIST
J
Junghun Yuk
KAIST
D
Dongyeon Kim
KAIST
Jihyo Kim
Jihyo Kim
KAIST InnoCORE PRISM-AI Center
KyungTae Lim
KyungTae Lim
École normale supérieure
Natural Language Processing