Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages

📅 2026-08-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the systemic neglect of low-resource languages in current AI infrastructure across data curation, tokenization, evaluation, and deployment, which exacerbates educational and linguistic inequities. Focusing on Bengali as a case study, the work integrates multilingual corpus analysis, tokenization efficiency benchmarks, internet penetration statistics, and modeling of educational resource accessibility to expose structural barriers: extreme training data scarcity (with an English-to-Bengali data ratio of 67:1), high tokenization overhead due to syllabic orthography, limited online content, and a pronounced rural–urban digital divide. The research reframes data scarcity not merely as a technical bottleneck but as a manifestation of structural injustice and advocates for an “offline-first” infrastructure design paradigm to advance linguistic equity and foster more inclusive AI development.
📝 Abstract
Artificial intelligence tools for education and language support are increasingly framed as scalable responses to access gaps in under-resourced communities. Yet the infrastructure underlying these tools, including training corpora, tokenization schemes, evaluation benchmarks, and deployment architectures, can systematically disadvantage speakers of underrepresented languages before a model is trained. This paper examines these structural barriers through Bengali, one of the world's most widely spoken languages, focusing on AI-assisted education in low-connectivity environments. We identify four interlocking failures: a severe web presence gap, with Bengali accounting for less than 0.5% of global web content despite representing nearly 4% of the global population; a 67:1 training-token deficit between English and Bengali in major multilingual corpora; a tokenization penalty associated with Bengali's alphasyllabary script that compounds the data deficit through higher token fertility; and connectivity exclusion, with individual internet penetration at 36.5% in rural areas compared with 71.4% in urban areas. These failures reflect longstanding resource-allocation decisions, institutional priorities, and design defaults that did not center underrepresented languages in mainstream AI development. We argue that dataset scarcity should be understood as a structural barrier rather than an isolated technical limitation, and that offline-first design should be treated as an equity-oriented infrastructure strategy. We conclude with directions for linguistics and AI research aimed at reducing these structural inequalities.
Problem

Research questions and friction points this paper is trying to address.

underrepresented languages
AI infrastructure
structural barriers
dataset scarcity
language equity
Innovation

Methods, ideas, or system contributions that make the work stand out.

structural inequality
underrepresented languages
tokenization penalty
offline-first design
AI infrastructure
💼 Related Jobs
No related jobs found.
A
Avijit Roy
John Jay College of Criminal Justice, CUNY
P
Proma Roy
The City College of New York, CUNY