Cross-lingual Offensive Language Detection: A Systematic Review of Datasets, Transfer Approaches and Challenges
This paper addresses the challenge of cross-lingual offensive language detection in social media. We systematically review 67 studies, with the first comprehensive focus on cross-lingual transfer learning (CLTL) methods for this task. Methodologically, we propose a holistic classification framework tailored to CLTL, innovatively categorizing approaches by “transfer object” into three paradigms: instance-level, feature-level, and parameter-level transfer. We construct and publicly release two structured resource tables, comprehensively cataloging multilingual pre-trained models, dictionary-based alignment techniques, zero-/few-shot transfer strategies, and adversarial training methods. Key challenges—including linguistic imbalance, annotation scarcity, and cultural context deficiency—are distilled and analyzed. The work yields a reusable research roadmap and an open resource repository, providing both theoretical foundations and practical tools for cross-lingual harmful content governance.