🤖 AI Summary
This study addresses the systemic neglect of low-resource languages in current AI infrastructure across data curation, tokenization, evaluation, and deployment, which exacerbates educational and linguistic inequities. Focusing on Bengali as a case study, the work integrates multilingual corpus analysis, tokenization efficiency benchmarks, internet penetration statistics, and modeling of educational resource accessibility to expose structural barriers: extreme training data scarcity (with an English-to-Bengali data ratio of 67:1), high tokenization overhead due to syllabic orthography, limited online content, and a pronounced rural–urban digital divide. The research reframes data scarcity not merely as a technical bottleneck but as a manifestation of structural injustice and advocates for an “offline-first” infrastructure design paradigm to advance linguistic equity and foster more inclusive AI development.
📝 Abstract
Artificial intelligence tools for education and language support are increasingly framed as scalable responses to access gaps in under-resourced communities. Yet the infrastructure underlying these tools, including training corpora, tokenization schemes, evaluation benchmarks, and deployment architectures, can systematically disadvantage speakers of underrepresented languages before a model is trained.
This paper examines these structural barriers through Bengali, one of the world's most widely spoken languages, focusing on AI-assisted education in low-connectivity environments. We identify four interlocking failures: a severe web presence gap, with Bengali accounting for less than 0.5% of global web content despite representing nearly 4% of the global population; a 67:1 training-token deficit between English and Bengali in major multilingual corpora; a tokenization penalty associated with Bengali's alphasyllabary script that compounds the data deficit through higher token fertility; and connectivity exclusion, with individual internet penetration at 36.5% in rural areas compared with 71.4% in urban areas.
These failures reflect longstanding resource-allocation decisions, institutional priorities, and design defaults that did not center underrepresented languages in mainstream AI development. We argue that dataset scarcity should be understood as a structural barrier rather than an isolated technical limitation, and that offline-first design should be treated as an equity-oriented infrastructure strategy. We conclude with directions for linguistics and AI research aimed at reducing these structural inequalities.