Latent Mechanisms of Language Control in Multilingual Language Models

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决多语言模型中不必要的代码转换问题,通过比较三种方法(ValSel、FreqSel和AnnSel)识别控制语言的潜在机制,实验表明这些方法有效。
📝 Abstract
Multilingual large language models can exhibit unintended code-switching -- unnecessarily alternating between languages during generation. We present a comparative study of three methods that identify language-controlling latents in cross-layer transcoders: activation value-based selection (ValSel), activation frequency-based selection (FreqSel), and LLM-generated latent annotation-based selection (AnnSel). To evaluate the efficacy of these methods in identifying language-controlling latents, we introduce two multilingual benchmarks that exhibit code-switching for fine-grained analysis of language steering across seven languages. Through targeted intervention experiments on Gemma-2-2B and Qwen3-4B, we find that all three methods effectively manipulate generation language, with FreqSel achieving the strongest overall performance, while AnnSel offering interpretable latent selection through explicit language annotations. A knock-out analysis suggests the methods select non-overlapping but each-functional latent subsets, indicating redundancy rather than a single canonical language direction. Code and data can be found at https://github.com/rm-3284/Latent-Mechanism-Multilingual.
Problem

Research questions and friction points this paper is trying to address.

Multilingual Language Models
Code-Switching
Language Control
Innovation

Methods, ideas, or system contributions that make the work stand out.

activation value-based selection
activation frequency-based selection
latent annotation-based selection
code-switching
multilingual benchmarks
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Ryo Mitsuhashi
Princeton University
S
Sabri Boughorbel
Prince Sattam bin Abdulaziz University
Majd Hawasly
Majd Hawasly
QCRI, Hamad Bin Khalifa University
Autonomous systemsLifelong learningNatural Language Processing