Listen to a native French speaker order coffee, or a child in rural Karnataka spin a story in Kannada. You aren’t just hearing words. You’re brushing up against a layered system shaped by centuries of migration, trade, conquest, and quiet domestic routines. Language isn’t a uniform code; it’s a living record of human experience. So why, when we build computational systems to navigate this terrain, do we expect them to treat every language with the same ease? The outcome, predictably, is a patchwork of brilliance and baffling failure.
Here’s the puzzle. A system parses a dense English legal document with sharp precision. Then it fumbles a simple Swahili news summary. The same architecture captures the weight of a Japanese haiku but trips over the politeness registers baked into Korean verb endings. The reasons aren’t mysterious, but they are tangled—rooted in data, linguistics, and the history of how we encode knowledge.
This unevenness isn’t a temporary glitch. It reflects fundamental asymmetries in the digital world and in the structures of the languages themselves. To get a grip on it, we need to look beneath benchmark scores and into the messy, glorious diversity of human speech.

The Ghost of Data Scarcity
The most immediate—and in many ways the most stubborn—source of uneven performance is the sheer disparity in available training data. Modern multilingual systems are voracious readers of text, and the internet, for all its global reach, is a lopsided archive. English, by many estimates, makes up roughly half of all web content. The so-called high-resource languages—Mandarin, Spanish, German, French, Japanese—follow, with deep digital footprints across news archives, parliamentary records, encyclopedias, and sprawling social platforms.
Now think about Oromo. It’s spoken by tens of millions in Ethiopia and Kenya, yet its presence in digitized, easily crawlable formats is thin. A system trained on a global web crawl will see English-language discussions of Oromo grammar far more often than actual Oromo. The model isn’t learning the language. It’s learning a linguistic shadow. This creates a grim feedback loop: low-resource languages lack the data to train capable systems, and the absence of capable systems means less incentive to create digital content in those languages.
Even within a single language, the data tilts. The text available online isn’t a neutral sample of human communication. It leans formal, written, and comes from a narrow demographic slice. A model trained on this will wrestle with the colloquial, the spoken, the dialectal. It might conjugate a textbook Yoruba verb flawlessly and then stare blankly at a lively, code-switched conversation in a Lagos market. The data isn’t just scarce. It’s biased toward a particular, sanitized register.
The Long Tail of Linguistic Resources
The problem stretches beyond raw text. Supervised datasets for specific tasks—like question answering, sentiment analysis, or textual entailment—cluster overwhelmingly in English and a few other languages. To build a sentiment analyzer for Urdu, you need thousands of examples where a human has labeled a sentence as positive, negative, or neutral. That work is expensive and slow. It demands native speakers and annotation guidelines that account for cultural context. A thumbs-up emoji might signal approval in one culture and deep offense in another. Without this task-specific data, a model’s general linguistic knowledge sits inert; it can generate fluent Urdu but can’t reliably tell you if a film review is a pan or a rave.
This is why benchmarking suites often show a dramatic resource cliff. A single model might hit near-human levels on an English reading comprehension test, then plummet to near-random chance on a comparable test in Amharic. The architecture is capable. The scaffolding of labeled examples just isn’t there to hold it up.

When Grammar Itself Becomes a Wall
Even with mountains of data, some languages throw up structural challenges that resist the assumptions baked into most architectures. The dominant paradigm leans on tokenization: breaking text into discrete units, usually words or sub-words. This approach, tuned for languages with clear whitespace boundaries like English, stumbles when it hits the fluidity of other writing systems.
Take Thai or Lao, where words are written without spaces. A string like การแบ่งคำในภาษาไทย isn’t one unbroken word. It’s a sequence of several words (“word segmentation in Thai language”). The model first has to infer where one word ends and the next begins—a task that itself requires deep linguistic knowledge. Errors in this initial segmentation cascade through every later layer of processing. A system might get tangled up with the character การ, which can be a standalone noun meaning “work,” but here is part of a longer nominalizing prefix. The boundary between vocabulary and grammar turns into a blurry, moving target.
Morphology: The Curse of Riches
If segmenting Thai is a puzzle of too few signals, morphologically rich languages present a problem of too many forms. An English verb like “walk” has four inflected forms: walk, walks, walked, walking. A Turkish verb like “okumak” (to read) can spawn thousands of forms, each a dense package of meaning. Okutamadıklarımızdan mısınız? translates to “Are you one of those whom we could not cause to read?”
In English, that meaning spreads across a dozen separate words, each a distinct token the model can learn independently. In Turkish, it’s fused into a single, complex word, built from a root plus a chain of suffixes for negation, ability, causation, tense, person, and a relative clause marker. A sub-word tokenizer will slice this word into smaller pieces—oku, ta, ma, dık, lar, ımız, dan, mı, sınız—but the model then has to learn that the meaning of the whole isn’t simply the sum of these often-ambiguous bits. The piece -lar is a plural marker, but -dan could be an ablative case marker (“from”). The combinatorial explosion of possible forms means the model encounters most words very rarely, a phenomenon known as data sparsity within a supposedly high-resource context.
Finnish, Hungarian, and many Indigenous languages of the Americas share this trait. A system trained mostly on the isolating patterns of English or Mandarin faces a steep climb. The model’s internal representations—its vectors for encoding meaning—grow more diffuse and less reliable as they try to capture the fluid, compositional nature of these words.
The Hidden Architecture of Semantics
The unevenness drills down past surface forms. It touches the very way meaning is organized. Languages aren’t just different word lists for the same catalog of concepts. They carve up the world along different joints. A model trained overwhelmingly on English absorbs a specific ontology: a way of categorizing objects, actions, and relationships.
Color terms are a classic example. English has “blue.” Russian distinguishes between light blue (голубой) and dark blue (синий). An English-centric model will tend to collapse this distinction, treating both as synonyms or near-synonyms. Asked to translate, it might render a sentence about a голубой sky as simply “blue,” dropping a piece of information that is fundamental and obligatory for a Russian speaker. The model isn’t just making a lexical error. It’s failing to perceive a conceptual boundary its primary training language taught it to ignore.
Spatial reasoning can diverge even more starkly. Many languages use egocentric coordinates (left, right, front, back) for spatial descriptions, like English does. But languages like Guugu Yimithirr, spoken by an Aboriginal community in Queensland, Australia, lean on cardinal directions (north, south, east, west) even for small-scale space. A speaker would say, “There’s an ant on your southeast leg.” A model soaked in English spatial frames has no native mechanism for this. It has to learn an entirely different computational logic for talking about where things are, a logic that demands constant, implicit geospatial tracking. The model’s struggle isn’t with vocabulary. It’s with a profoundly different way of anchoring thought to the physical world.

Pragmatics and the Unspoken Rules
Beyond semantics lies the minefield of pragmatics: the social rules that govern language use. This is where politeness, formality, and indirectness live. Japanese and Korean have complex honorific systems that are grammatically encoded. A verb form doesn’t just describe an action; it signals the relative social status of the speaker, the listener, and the person being discussed. Picking the wrong level—masu vs. desu, or the even more layered Korean jondaemal vs. banmal—can be a deep social gaffe, turning a polite inquiry into an insult.
Models trained on vast, unsorted text corpora wrestle mightily with this. They can learn the forms, but they lack a grounded feel for the social context that dictates their use. The data itself may be a mess, mixing formal and informal registers from different sources. The result is often a strangely oscillating tone: a single paragraph might lurch from stiff, textbook politeness to casual intimacy, creating a textual persona that feels socially unmoored and, to a native speaker, deeply unsettling. The model fails not at grammar, but at being a socially competent member of the speech community.
The Script as a Cognitive Barrier
We tend to think of a writing system as a mere container for language, but it can be a significant source of computational bias. Models that lean on a shared sub-word vocabulary for multiple languages have to make hard choices about what to include. Scripts with large character inventories, like Chinese, Japanese, and Korean (CJK), compete for space in the model’s tokenizer with the relatively small Latin alphabet used by English and many others.
This competition is rarely fair. A tokenizer optimized for byte-pair encoding across 100 languages will naturally allocate more slots to the frequent byte pairs found in the dominant high-resource scripts. That can lead to a phenomenon called over-tokenization for other scripts. A single Chinese character, rich in semantic and phonetic information, might get split into multiple meaningless byte-level tokens. A word in Devanagari, the script used for Hindi, might be broken into pieces that don’t line up with its natural syllabic structure. The model is forced to learn meaning from shattered fragments, a task that requires more data and more computation to overcome. The very first stage of processing, before any meaning is extracted, imposes an unfair tax on non-Latin scripts.
Toward a More Genuine Multilingualism
The uneven performance of multilingual models isn’t a sign of failure. It’s a map of our own technological and cultural priorities. It reflects the digital divide, the hard problem of linguistic structure, and the deep challenge of encoding cultural pragmatics. Noticing this unevenness is the first step toward addressing it. The goal shouldn’t be a bland, uniform performance that erases linguistic difference. It should be a set of tools that is honest about its limits and respectful of the unique character of each language it tries to learn.
Progress won’t come from a single architectural breakthrough. It lies in a more ecologically minded approach: curating better, more representative datasets for low-resource languages; designing tokenizers that are script-aware and morphologically sensitive; and developing evaluation benchmarks that probe for genuine understanding of pragmatics and cultural context, not just surface-level pattern matching. The most eloquent polyglot isn’t the one who speaks every language identically. It’s the one who knows what they don’t know and listens for the deeper structures beneath the words.
Frequently Asked Questions
Why do some languages require vastly more data to train a model well?
This often comes down to a mix of morphological complexity and script-related headaches. A language like Turkish or Finnish, which packs a great deal of meaning into single, highly variable words, presents a much larger vocabulary for a model to learn. The model encounters most word forms very rarely, making it harder to learn solid patterns. At the same time, scripts without clear word boundaries, like Thai, or large-character scripts like Chinese, can be handled inefficiently by standard tokenizers, effectively wasting the model’s capacity on basic segmentation tasks rather than higher-level understanding.
Does a model’s trouble with politeness levels in Japanese mean it doesn’t “understand” the language?
It understands a version of the language, but one flattened by its training data. A model can learn the grammatical forms of honorifics, but it learns them from a mixed corpus of novels, news, and casual web text. It lacks a consistent model of the social world that dictates when to use them. The result is often grammatically correct but socially incoherent output, revealing that understanding for a computational system is built on statistical correlations in text, not on the grounded, lived experience of social interaction that humans use to navigate linguistic register.
Can we fix the imbalance in digital language resources?
The imbalance is a structural feature of the current internet, and it won’t be fixed overnight. But a more productive path than simply scraping more web data is the careful, community-driven creation of targeted datasets. This includes not just raw text, but parallel corpora for translation, supervised datasets for specific tasks like question answering, and resources that capture spoken, colloquial, and dialectal varieties. These efforts are slow and require deep collaboration with native speakers, but they build the necessary scaffolding for models to move beyond a high-resource echo chamber and into genuine linguistic diversity.
The uneven polyglot, in its current form, is a mirror. It shows us not a failing of technology, but a reflection of our own unevenly connected, richly diverse, and deeply complex linguistic world. The task ahead isn’t to smooth out the reflection. It’s to build a lens that can appreciate the full depth of the field.