AI Summary • Published on Aug 11, 2026
Artificial intelligence tools, especially those for education, are frequently promoted as scalable solutions to address educational disparities in under-resourced communities. However, this paper argues that the foundational infrastructure underpinning these AI tools—including their training datasets, tokenization methods, evaluation benchmarks, and deployment architectures—is built upon assumptions that inherently disadvantage speakers of underrepresented languages. The authors use Bengali, a language with hundreds of millions of speakers globally, as a detailed case study to illustrate how these design choices create significant barriers for learners, particularly in environments with limited internet access. This systematic exclusion, termed "structural silence," stems not from explicit policies but from the cumulative impact of design decisions that did not prioritize these languages during AI development.
This research is presented as an analytical synthesis and a case study, aiming to provide a comprehensive explanation rather than introducing new empirical data or system contributions. The primary objective is to consolidate existing technical, infrastructural, and educational evidence to clarify why AI support remains uneven for Bengali-speaking learners and to identify the underlying structural logic perpetuating this disparity. The paper proceeds by first reviewing the relevant background in low-resource language Natural Language Processing (NLP) and AI infrastructure. It then meticulously identifies and supports with evidence four distinct structural failures. Subsequently, these failures are linked to their cognitive and educational consequences for Bengali-speaking learners. The analysis draws upon established educational theories, such as Cognitive Load Theory, alongside published performance benchmarks and public infrastructure indicators related to web presence, data availability, and internet connectivity.
The paper identifies four interconnected structural failures that collectively disadvantage speakers of underrepresented languages within AI infrastructure. First, there is a significant **web presence gap**: despite Bengali speakers constituting approximately 4% of the global population, the language accounts for less than 0.5% of global web content. This stark disparity directly translates into a severe **training token deficit**, where major training resources show an approximate 67:1 ratio of English to Bengali tokens. Third, a **tokenization penalty** arises because standard tokenizers, optimized for Latin-script languages, process Bengali’s alphasyllabary script inefficiently, requiring significantly more subword tokens (higher token fertility) to represent the same semantic content. This fragmentation increases computational overhead and disrupts linguistic units, degrading model performance. Finally, a **connectivity exclusion** makes cloud-dependent AI tools functionally inaccessible to populations in low-connectivity rural areas, where individual internet penetration can be less than half that of urban areas. These failures culminate in compounded cognitive demands for Bengali-speaking learners using AI tools, especially when explanations are in English, hindering comprehension and learning outcomes.
The findings suggest that the scarcity of datasets for low-resource languages should be understood as a structural barrier, stemming from historical resource allocation and institutional priorities, rather than merely a technical challenge. The paper proposes three critical responses for the linguistics and AI communities. Firstly, work on low-resource language infrastructure, including the creation of datasets, benchmarks, and evaluation protocols, must be recognized as primary research contributions, rather than secondary or preliminary efforts. Secondly, designing AI systems with **offline-first capabilities** should be embraced as a key equity-oriented strategy, not just a technical compromise, to extend accessibility to communities with unreliable internet access. Such designs also offer sustainability benefits. Thirdly, **linguistic analysis** is crucial for identifying and articulating the underlying, often invisible, language-specific assumptions embedded within AI infrastructure (e.g., in tokenization schemes) that currently disadvantage non-English languages. Addressing these systemic issues requires sustained institutional commitment and interdisciplinary collaboration to foster equitable AI development.