Code-Mix Sentiment Analysis on Hinglish Tweets

📅 2026-01-08
🏛️ arXiv.org
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges posed by Hinglish—a prevalent Romanized Hindi–English code-mixed variety on Indian social media—whose spelling variations, slang, and out-of-vocabulary terms significantly degrade the performance of conventional monolingual NLP models in sentiment analysis, thereby undermining brand monitoring efforts. To tackle this, the work proposes a high-performance sentiment classification framework specifically designed for Hinglish tweets, which uniquely integrates fine-tuned multilingual BERT (mBERT) with subword tokenization to effectively handle the low-resource nature of code-mixed language. Evaluated on public benchmarks, the approach achieves state-of-the-art accuracy, offering both a deployable, production-ready tool for brand sentiment tracking and a new strong baseline for code-mixing NLP tasks.

Technology Category

Application Category

📝 Abstract
The effectiveness of brand monitoring in India is increasingly challenged by the rise of Hinglish--a hybrid of Hindi and English--used widely in user-generated content on platforms like Twitter. Traditional Natural Language Processing (NLP) models, built for monolingual data, often fail to interpret the syntactic and semantic complexity of this code-mixed language, resulting in inaccurate sentiment analysis and misleading market insights. To address this gap, we propose a high-performance sentiment classification framework specifically designed for Hinglish tweets. Our approach fine-tunes mBERT (Multilingual BERT), leveraging its multilingual capabilities to better understand the linguistic diversity of Indian social media. A key component of our methodology is the use of subword tokenization, which enables the model to effectively manage spelling variations, slang, and out-of-vocabulary terms common in Romanized Hinglish. This research delivers a production-ready AI solution for brand sentiment tracking and establishes a strong benchmark for multilingual NLP in low-resource, code-mixed environments.
Problem

Research questions and friction points this paper is trying to address.

Code-Mix
Sentiment Analysis
Hinglish
NLP
Multilingual
Innovation

Methods, ideas, or system contributions that make the work stand out.

code-mixing
Hinglish
mBERT
subword tokenization
sentiment analysis
🔎 Similar Papers
No similar papers found.
A
Aashi Garg
Indira Gandhi Delhi Technical University for Women, Kashmere Gate, Delhi
A
Aneshya Das
Indira Gandhi Delhi Technical University for Women, Kashmere Gate, Delhi
A
Arshi Arya
Indira Gandhi Delhi Technical University for Women, Kashmere Gate, Delhi
A
Anushka Goyal
Indira Gandhi Delhi Technical University for Women, Kashmere Gate, Delhi
A
Aditi
Indira Gandhi Delhi Technical University for Women, Kashmere Gate, Delhi