Detecting Linguistic Diversity on Social Media

📅 2025-02-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of New Zealand’s official language data—reliant solely on decennial censuses and lacking fine-grained, dynamic monitoring capabilities—by proposing a novel paradigm for assessing linguistic diversity using social media data (the CGUL Twitter corpus). Methodologically, it integrates multi-model language identification, geospatial matching, and cross-source data alignment to compute linguistic diversity indices at national, regional, and local administrative levels, systematically benchmarked against census benchmarks. Empirically, it demonstrates for the first time that social media–derived language data exhibit high sensitivity and reliability in capturing subnational language ecology and sociolinguistic change: significant correlations (p < 0.01) are observed between social media–based estimates and census figures at regional and local scales. The findings establish the feasibility of leveraging social media as a complementary, near–real-time language monitoring tool, thereby providing a new empirical foundation for evidence-based language policy and minority language preservation.

Technology Category

Application Category

📝 Abstract
This chapter explores the efficacy of using social media data to examine changing linguistic behaviour of a place. We focus our investigation on Aotearoa New Zealand where official statistics from the census is the only source of language use data. We use published census data as the ground truth and the social media sub-corpus from the Corpus of Global Language Use as our alternative data source. We use place as the common denominator between the two data sources. We identify the language conditions of each tweet in the social media data set and validated our results with two language identification models. We then compare levels of linguistic diversity at national, regional, and local geographies. The results suggest that social media language data has the possibility to provide a rich source of spatial and temporal insights on the linguistic profile of a place. We show that social media is sensitive to demographic and sociopolitical changes within a language and at low-level regional and local geographies.
Problem

Research questions and friction points this paper is trying to address.

Assessing linguistic diversity using social media data.
Comparing social media and census data for language use.
Analyzing linguistic changes at national, regional, and local levels.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Social media data for linguistic diversity analysis
Comparison with census data as ground truth
Language identification models for validation