IndicTriMix: Developing Language Identification Datasets and Models for Tri-Language Code-Mixing

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究针对社交媒体中的三语代码混用现象,通过序列标注问题和微调基于上下文的转换模型来准确识别代码混用文本中的语言。
📝 Abstract
Language identification in code-mixed text, largely observed in social media, is highly essential when users frequently switch between multiple languages within a single utterance. Accurately identifying the languages of code-mixed tokens becomes an urgent necessity. Traditional language identification models, designed for monolingual text, are not well suited for token-level language identification in code-mixed settings. We formulate the task as a sequence labeling problem and fine-tune contextual transformer-based models MuRIL and XLM-RoBERTa best suited for Indian languages. We evaluate these systems on three different data configurations (Hindi, Gujarati, and Bengali) to predict language labels for individual tokens. We release a benchmark for language identification in code-mixed tokens with manually annotated test sets. We propose two approaches of code-mixed generation using parallel sentences of three languages. The trained models demonstrate the effectiveness of contextual embeddings for token-level language identification in multilingual social media text. For reproducibility and to facilitate future research, we publicly release our fine-tuned models.
Problem

Research questions and friction points this paper is trying to address.

language identification
code-mixed text
social media
multilingual
Innovation

Methods, ideas, or system contributions that make the work stand out.

code-mixed text
sequence labeling
contextual embeddings
multilingual social media
🔎 Similar Papers
No similar papers found.