PERCEPT: A Corpus for POS Tagging and Analysis of Persian-English Code-Mixing

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of a Persian–English code-switched corpus annotated with Universal Dependencies (UD) part-of-speech tags, a gap that has hindered linguistic analysis and the development of syntax-aware NLP models for such mixed-language data. To bridge this gap, we introduce PERCEPT, the first large-scale, multi-platform code-switched corpus comprising 6,800 posts collected from X, Instagram, and Digikala. We further propose a large language model–assisted framework for automatic UD annotation, delivering the first publicly available UD-compliant POS tags for Persian–English code-switched text. Human evaluation confirms high agreement between our automatically generated annotations and gold-standard labels. Linguistic analysis reveals that nouns dominate as the primary category involved in code-switching, the positional distribution of switches remains consistent across platforms, and language-triggering effects are markedly stronger in Digikala.
📝 Abstract
Social media has become a major venue for multilingual communication, where users frequently mix multiple languages within a single utterance. Although code-mixed corpora have been developed for several language pairs, Persian-English code-mixing remains relatively underexplored. Existing Persian resources lack Universal Dependencies (UD) part-of-speech (POS) annotations for code-mixed words, limiting both linguistic analyses and the development of syntax-aware NLP models. To address this gap, we introduce PERCEPT, the first publicly available large-scale Persian-English code-mixed corpus annotated with Universal Dependencies POS tags for code-mixed words. The dataset comprises 6,800 posts collected from X, Instagram, and Digikala. We further present an LLM-assisted annotation framework that automatically assigns POS tags and document-level topics. Human evaluation demonstrates high agreement between the automatically generated annotations and gold annotations, confirming the reliability of the annotations. Using PERCEPT, we conduct the first comprehensive linguistic analysis of Persian-English code-mixing across multiple social media platforms. Our analyses reveal that nouns are the predominant category for code-mixed words, while the distributions of other POS categories vary across platforms. We further find that the positional distribution of code-mixed words is remarkably consistent across platforms, whereas the triggering effect is substantially more pronounced in Digikala. PERCEPT is publicly available at https://github.com/kalhorghazal/PERCEPT.
Problem

Research questions and friction points this paper is trying to address.

code-mixing
Persian-English
POS tagging
Universal Dependencies
corpus
Innovation

Methods, ideas, or system contributions that make the work stand out.

code-mixing
Universal Dependencies
POS tagging
LLM-assisted annotation
Persian-English corpus
G
Ghazal Kalhor
School of Electrical and Computer Engineering, College of Engineering, University of Tehran, Tehran, Iran
Z
Zahra Jafari
School of Engineering Science, College of Engineering, University of Tehran, Tehran, Iran
A
Amirarsalan Shahbazi
School of Electrical and Computer Engineering, College of Engineering, University of Tehran, Tehran, Iran
Behnam Bahrak
Behnam Bahrak
Tehran Institute for Advanced Studies