Walk into a library in Tokyo, and you’ll see a system that instantly catalogs a Japanese book, catching the fine line between a title and an author’s name. Fly to a library in Nairobi, and that same underlying software might trip over a Swahili headline, mistaking a common noun for a proper name. It’s not a bug. It’s a structural imbalance—a reflection of how computational systems learn to parse human language based on historical data, typological complexity, and plain old economic priority.

The Resource Curse of High-Resource Languages
When we talk about a model’s “fluency,” we’re really talking about statistical density. English dominates digitized text—not just because of its native speakers, but because the early internet, academic publishing, and open-source code repositories were overwhelmingly Anglophone. A model trained on a snapshot of the web will absorb petabytes of meticulously labeled English data. For Hindi, a language with hundreds of millions of speakers, the available clean, digitized corpus is orders of magnitude smaller. The model isn’t inherently less capable; it’s starved of nourishment.
This imbalance creates a vicious cycle. Developers optimize tokenizers for English morphology. A tokenizer that neatly splits “unbelievable” into “un-believ-able” will butcher a Turkish word like “evlerinizden” (from your houses) into a dozen meaningless fragments. The model then has to piece together meaning across a much longer sequence, which is computationally expensive and error-prone. The very tools meant to smooth language processing end up penalizing languages that don’t fit the English mold.
Morphological Mazes and Syntactic Snags
Linguistic structure itself is a major hurdle. English is relatively analytic, leaning on word order and helper words. Finnish or Hungarian, by contrast, are highly synthetic, cramming immense grammatical detail into suffixes. A model raised on English patterns struggles to adapt to a system where a single word can carry the load of an entire English phrase. The internal representations—vectors, attention patterns—learned from English don’t map neatly onto these different grammatical architectures.

Take grammatical gender. In English, gender is mostly natural—he, she, it—applying to animate beings. In German or Russian, every noun has a grammatical gender that affects articles, adjectives, and pronouns. A model trained on English has no native mechanism to track this. It has to learn gender agreement as a secondary, often brittle pattern. The result? A higher error rate in generating coherent sentences, where an adjective might modify a noun of the wrong gender—a mistake that instantly marks the output as non-native.
The Script Barrier
Beyond grammar, the very symbols used to write a language can cause uneven performance. Models typically tokenize text, breaking it into smaller units. For alphabetic scripts like Latin or Cyrillic, this is fairly efficient. For logographic scripts like Chinese, where each character can represent a morpheme or a whole word, tokenization strategies often fall apart. A common workaround is to treat each character as a token, but that explodes sequence length and dilutes semantic density. The model sees a stream of characters without the word boundaries that spaces provide in English, making it harder to learn meaningful groupings.
Languages using the Arabic script add another layer of difficulty. The cursive nature of the script—letters change shape depending on their position—combined with the frequent omission of short vowels in writing, creates a high degree of ambiguity. A single written form can correspond to several different spoken words, differentiated only by vowel markings that are often absent. A model must resolve this ambiguity from context, a task that demands a much deeper, more layered understanding of the text than is needed for a fully vocalized script like Spanish.
Data Quality, Not Just Quantity
Even when data exists for a language, its quality can silently kill performance. Web-scraped corpora for many languages are full of code-switching—the mixing of two or more languages in a single sentence or document. In many post-colonial nations, formal documents may be in English or French, while informal online discourse blends the official language with local tongues. A model trained on this noisy data learns a distorted representation of the “pure” language, often defaulting to the dominant language’s grammar when it encounters ambiguity.
Translationese is another subtle contaminant. A significant portion of multilingual data comes from translated texts, especially for mid-resource languages. Translated text, even when done by expert humans, carries the stylistic and structural fingerprints of the source language. Models trained on this data learn a version of, say, Swahili that is syntactically calqued from English. They produce output that is grammatically correct but stylistically foreign, lacking the idiomatic rhythm of naturally authored prose.

The Dialect Continuum and Standardization
The concept of a “language” itself is a political and social construct that models struggle to navigate. Arabic is a classic example. Modern Standard Arabic (MSA) is used in formal writing across the Arab world, but no one speaks it at home. The spoken varieties—Moroccan Darija, Egyptian Ammiya, Levantine—are so divergent that they can be mutually unintelligible. A model trained on a corpus labeled simply “Arabic” is usually trained overwhelmingly on MSA. When presented with a tweet in Egyptian dialect, its performance plummets. It’s being asked to process a language it has barely seen, yet both are filed under the same label.
This standardization bias also affects languages with deep orthographic variation. Norwegian, with its two official written standards (Bokmål and Nynorsk), or Serbian, which officially uses both Cyrillic and Latin scripts, forces models to split their already limited resources. The model doesn’t see one language; it sees two or three closely related but distinct data streams, each too small to build a strong representation on its own.
Cultural Context and Named Entity Recognition
A particularly stark measure of unevenness is Named Entity Recognition (NER)—the task of identifying names of people, places, and organizations in text. A model trained on English news excels at spotting “John Smith” and “New York City.” Apply it to a Thai news article, and it fails to recognize “กรุงเทพมหานคร” (Bangkok) because the entity’s romanized form, which might appear in the training data, bears no orthographic resemblance to its native script. The model’s world knowledge is fundamentally tied to the script and cultural context of its primary training data.
This extends to cultural references. A model can effortlessly link “Barack Obama” to a dense web of associated concepts. It has no such rich representation for “Jomo Kenyatta” or “Kim Dae-jung,” not because they are less historically significant, but because the digital footprint of their names in English-language training data is smaller. The model’s “intelligence” is a mirror of the internet’s Anglophone bias, reflecting a distorted map of human knowledge where some regions are rendered in high definition and others are barely sketched.
Benchmarks That Don’t Travel
Our methods for measuring performance are themselves part of the problem. Standard benchmarks are often translated from English. A reasoning task designed to test logical deduction in English is translated word-for-word into Vietnamese. But the translation process can introduce artifacts—unnatural phrasing, calqued idioms—that make the task harder for the model for reasons unrelated to reasoning ability. The model isn’t failing at logic; it’s failing to parse a stilted, unnatural sentence that no native speaker would ever produce.
Additionally, benchmarks rarely capture the functional needs of non-English speakers. A model might score highly on a sentiment analysis test for Spanish movie reviews but fail miserably at extracting the correct dosage from a medical prescription written in Tagalog. The tasks that matter for daily life in different linguistic communities are not equally represented in the evaluation suites that drive research and development. Performance is optimized for the tests we have, not the tasks people need.
The Path Toward More Even Ground
Addressing this unevenness requires a shift in focus from model architecture to data ecology. For low-resource languages, the most impactful work is not tweaking a transformer layer but building and curating high-quality, naturally authored corpora. This is slow, unglamorous work that involves digitizing local newspapers, recording and transcribing oral histories, and creating open-source datasets that reflect the true idiomatic character of a language, not its translated shadow.
Techniques like targeted pre-training can help. A model first trained on a massive multilingual corpus can undergo a second phase of intensive training on a smaller, high-quality dataset for a specific language. This allows the model to retain the general syntactic intelligence learned from the larger pool while adapting its representations to the specific morphology and vocabulary of the target language. It’s a form of academic exchange: the model brings general knowledge, then immerses itself in a local linguistic environment to gain native-like fluency.
Script and tokenization issues demand bespoke solutions. For languages with complex morphology, using morpheme-level tokenization instead of standard sub-word algorithms can yield dramatic improvements. For logographic scripts, incorporating visual or radical-based features alongside textual ones can help the model ground characters in their constituent parts, much as a human learner uses radicals to guess the meaning and pronunciation of an unknown Chinese character.
Ultimately, the uneven performance of multilingual models is a cartographic problem. We have mapped the linguistic world using a projection that centers on English and distorts everything else. Correcting this distortion requires not just more data, but different kinds of data, collected and curated with a deep respect for the structural integrity of each language. The goal is not a single, monolithic model that speaks all languages with equal fluency—that may be a mathematical impossibility given the diversity of human grammar—but a constellation of systems, each deeply rooted in its own linguistic soil, yet capable of reaching across the divide.
Frequently Asked Questions
Why do some languages with many speakers, like Hindi or Bengali, still perform poorly in these systems?
Performance is tied not to the number of speakers, but to the volume and quality of digitized, machine-readable text available for training. Many languages with large speaker populations have a limited digital footprint because online communication often occurs in a dominant second language like English, or because local publishing industries have not fully transitioned to digital formats that are easily scraped and cleaned for training purposes.
What is code-switching and why does it affect model performance?
Code-switching is the practice of alternating between two or more languages within a single conversation or sentence. It is a natural and sophisticated linguistic behavior in many multilingual communities. However, when a model is trained on text that contains unsystematic code-switching, it can learn to blur the boundaries between languages. This leads to output that unpredictably mixes languages or applies the grammar of one language to the vocabulary of another, reducing clarity and coherence.
Can a model trained on translated text ever achieve native-like fluency?
It is very difficult. Translated text, even of high quality, carries structural echoes of the source language. A model trained predominantly on such data learns a version of the target language that is syntactically and stylistically influenced by the source. It may produce grammatically correct sentences that nonetheless feel unnatural to a native speaker because they lack authentic idiomatic expressions and follow the rhetorical patterns of a foreign tongue.