Lingua Franca or Probing Artifact? Rethinking Latent Language in Multilingual LLMs

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了多语言模型中潜在语言识别方法的差异,通过比较基于几何和解码的方法,揭示了它们在跨语言信息处理上显示的不同方面。
📝 Abstract
Latent language identification is often used to argue that multilingual language models route computation through language-specific states, such as English pivots. However, existing probes infer latent language from different signals, such as the geometry of hidden states or what can be decoded from intermediate representations. Since such claims shape conclusions about how models share and route information across languages, we ask whether these probes measure the same phenomenon or expose distinct aspects of multilingual computation. We study this question across model families, training regimes, domains, tasks, checkpoints, and up to 27 languages. We find that identification probes systematically disagree: the GMM-based representation probe, which draws evidence from hidden state geometry, shows earlier cross-lingual mixing, whereas decoding-based probes, which rely on output-space decodability, retain sharper language-specific and more English-biased signals. These differences track model multilinguality and training progression, but are comparatively stable across domains. Our results suggest a more cautious interpretation of latent language identification, where current probes expose different aspects of multilingual processing, rather than directly revealing a single internal lingua franca.
Problem

Research questions and friction points this paper is trying to address.

latent language
multilingual LLMs
probes
cross-lingual mixing
hidden state geometry
Innovation

Methods, ideas, or system contributions that make the work stand out.

latent language identification
multilingual LLMs
probing artifact
representation probe
decoding-based probes