Improving Language Identification for Code-Switched Utterances with Integer Linear Programming

📅 2026-09-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过改进MaskLID算法并引入整数线性规划,解决了代码转换语句在语言识别中的不足,提高了多语言组合识别的准确性和可解释性。
📝 Abstract
Automatic identification of code-switched (CS) utterances remains a challenge for language identification (LID) systems, causing such texts to be underrepresented in the training data of Large Language Models. In this paper, we revisit MaskLID, a state-of-the art approach for CS identification, which requires no training and detects arbitrary language combinations. We make three main contributions: (a) we reveal, and address, a major issue of MaskLID: its overreliance on word-level language association scores; (b) we reformulate the underlying optimization algorithm as an Integer Linear Program, enabling us to experiment with a large set of clear and interpretable constraints; (c) each of these improvements vastly improves the baseline system, as we illustrate in experiments involving 10~diverse languages, where we observe a strong boost in performance on CS benchmarks. We release our code and data for reproducibility.
Problem

Research questions and friction points this paper is trying to address.

Code-Switched
Language Identification
Training Data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Integer Linear Programming
Language Identification
Code-Switched Utterances
MaskLID
🔎 Similar Papers
No similar papers found.
J
Joanna Radoła
Sorbonne Université, CNRS, ISIR, Paris, France
J
Josep Maria Crego
SYSTRAN by ChapsVision, Paris, France
François Yvon
François Yvon
ISIR / CNRS et Sorbonne Université
Natural Language ProcessingSpeech ProcessingComputational LinguisticsMachine Translation