I still remember the first time the paradox really hit me. I was running a series of tests on a multilingual system—part of my computational linguistics work—and watched it breeze through a dense English financial report. Then, with the same architecture, it tripped over a simple Swahili children’s tale. Not just a small stumble. It mangled basic verb chains and lost track of who was doing what by the third sentence. That gap isn’t some random software hiccup. It’s a fingerprint left by deep structural rifts, old cultural grooves, and the uneven way we’ve built the digital record. Chasing that question pulled me into a strange borderland where linear algebra bumps up against ethnography, and it quietly rearranged how I think about language itself.

To see why multilingual models wobble so much, you have to look at the dirt they grew in: the training data. The web looks like a noisy global conversation, but the truth is lopsided. English alone swallows a ridiculous share of indexed text. Thousands of other languages barely register—a few scattered forum threads, some old scanned documents, maybe a handful of social media posts. When a model marinates in an ocean of English and gets only a thimbleful of Yoruba or Quechua, its internal map warps. It learns to track the fine grain of English syntax and implication with real delicacy. For a low-resource language, though, it’s working from shadows, extrapolating patterns it never actually observed. The result is a strange phantom competence: output that looks plausible from a distance but crumbles the moment you lean on it.
The Scarcity Trap: Data Asymmetries
Scarcity isn’t just about a small stack of words. It’s about missing texture. High-resource languages spill across registers—court opinions, bar arguments, lab reports, love letters, stand-up comedy transcripts. A model fed that sprawl learns to flex. Low-resource languages, by contrast, often show up in tight, repetitive strips: religious texts, bureaucratic forms, short social snippets. The model grows up in a linguistic monoculture. Ask it to write a playful story in Amharic, and it will likely default to the stiff cadence of the parliamentary records it was weaned on. That register poverty creates a fake sense of fluency. The grammar checks out, but the pragmatics—the living, breathing part—wobbles badly.

The trouble reaches deeper than style. Lexical gaps yawn open. A model might have seen “quantum entanglement” thousands of times in English, bracketed by rich context and clarifying sentences. For the same idea in Tamil, it might have glimpsed a single fuzzy phrase in a translated textbook, with almost no surrounding explanation. When you push for precision, the model isn’t pulling up a stored concept. It’s hastily stitching an approximation out of unrelated scraps, and the seam tears under even gentle pressure.
Tokenization and the Tyranny of Byte-Pair Encoding
Below the surface of words, there’s a mechanical snag that doesn’t get enough attention: tokenization. Most multilingual setups lean on subword methods like Byte-Pair Encoding, which build a vocabulary from frequent character chunks. It’s efficient, sure, but it carries a quiet bias. In a mixed-language model, the tokenizer greedily optimizes for whatever language dominates the data.
Take Thai or Lao. Words run together without spaces. A tokenizer raised mostly on English will carve up these scripts awkwardly—sometimes splitting a single word into a trail of meaningless fragments, sometimes gluing neighboring words into an unreadable blob. The model then has to learn meaning from a signal that’s already been mangled. A native speaker interacting with the system gets back a weird, splintered output. Not because the ideas are wrong, but because the system never really learned where one thought ends and another begins.
Morphological Complexity and the Agglutinative Challenge
Languages pack their grammar into wildly different containers. English, with its relatively bare inflections, is a gentle terrain. But Turkish, Finnish, or Inuktitut? They’re agglutinative—layer after layer of suffixes glued on to express what English needs a whole phrase to say. A single Turkish verb like “yapabileceklerdi” unpacks into “they would be able to do,” a tight little bundle of meaning.
For a model, that morphological richness is a combinatorial headache. The tokenizer, built for English-like brevity, often treats rare inflected forms as weird outliers. The model is forced to reason from fragments—“yap” + “abil” + “ecek” + “ler” + “di”—and hope the composition lands correctly. When the training data is thin, those compositions grow brittle. The model might handle a common form without issue but hallucinate wildly on a legitimate, less frequent inflection, spitting out something that means “they will be able to have been doing” when the intended sense was a straightforward past intention.

Isolating languages like Vietnamese throw a different kind of curveball. Without inflection to mark grammatical roles, meaning hangs heavily on word order and small function particles. A model trained on English’s explicit signals—the “-ed” for past, the “-s” for plural—has to learn to watch tiny words like “đã” or “đang” with extreme care. In low-data settings, those small, easy-to-overlook tokens turn invisible, and the result is a fog of temporal confusion in translated or generated text.
The Semantic Shadow of Culture
Language doesn’t float free of the world. Words are carriers of cultural weight, and translations are hardly ever neat one-to-one swaps. The Japanese term “和” (wa), often glossed as “harmony,” carries layers of social obligation, aesthetic judgment, and historical memory that no single English word can hold. When a multilingual model bumps into “wa,” it might anchor it to nearby English terms—“peace,” “balance,” “group”—but the quiet intensity, the sense that breaking wa is a deep social wound, stays unlearned.
This semantic slippage is everywhere. A model steeped in English-centric data will tend to link “individualism” with positive traits like freedom and creativity, absorbing a Western philosophical frame. Apply it to Korean, where collective identity is woven into the grammar itself (through honorifics, inclusive versus exclusive “we”), and the model can misread the whole tone, turning a neutral Korean statement into an English sentence that sounds oddly self-promotional or socially off-key. The unevenness here isn’t a translation error. It’s a failure of cultural translation.
Script and Sight: The Visual Dimension of Language
For plenty of languages, the first hurdle is simply being seen clearly. The Unicode standard covers a dazzling spread of scripts, but a model’s ability to handle them depends entirely on what it saw during training. Scripts like Devanagari (Hindi, Marathi, Nepali) or Ge’ez (Amharic, Tigrinya) often appear in limited, stylized corpora. The model learns to recognize the symbols but hasn’t seen enough variation in fonts, handwriting, or noisy OCR scans to build a solid visual intuition.
The practical result is fragility. A Hindi sentence typed in a slightly unusual font, or one that includes a rare conjunct consonant, can throw processing off the rails completely. The model doesn’t read the text. It stares at a pattern of shapes it can’t confidently decode. This gets worse for languages that are primarily spoken, where written standards are still fluid or disputed. A model trained on one dialect’s spelling may choke on another’s, fracturing a community that sees itself as unified.
The Echo of Historical Bias
Data isn’t neutral. It’s a fossilized record of human decisions. Colonial languages—English, French, Spanish—are wildly overrepresented because of old patterns of education, publishing, and technology access. Indigenous languages of the Americas or Australia, many now in active revitalization, remain digitally marginal. When a model is trained on this skewed archive, it quietly reinforces a hierarchy. A query in Welsh might pull back an answer that leans on English Wikipedia’s framing of Welsh history, gently overwriting local perspectives without anyone noticing.
Even inside well-resourced languages, dialect variation gets short shrift. A model trained on Standard Arabic will stumble over Moroccan Darija, which mingles Arabic, Berber, and French into a vernacular with no official written standard. The system might treat Darija as “noisy Arabic,” trying to correct it into a formal register that washes away the speaker’s identity. That’s not a simple technical glitch. It’s a kind of erasure, baked right into the code.
Evaluating the Unevenness
So how do we even measure this unevenness? The usual benchmarks are part of the trouble. A translation quality metric like BLEU leans on reference translations that tend to exist mainly in high-resource languages. For a low-resource language, those reference texts may be scarce, poorly checked, or produced by non-native speakers. A model can post a strong score on a bent yardstick while generating output that native speakers find stiff, odd, or outright offensive. We’re measuring shadows and calling it precision.
Human evaluation, the supposed gold standard, brings its own distortions. Evaluators fluent in, say, both English and Swahili are rare. When we do find them, their judgments often carry the imprint of their own schooling—frequently in English-dominant institutions. They might unconsciously penalize Swahili outputs that echo oral, poetic traditions because those don’t match the formal written style they were taught to respect. The unevenness gets reinforced by the very tools meant to spot it.
Toward More Equitable Multilingual Systems
Untangling these disparities asks for a shift in posture. The aim can’t be to cram every language into an English-shaped container. We need architectures that respect linguistic variety from the ground up. That could mean language-specific tokenizers that actually understand the script, or training objectives that explicitly reward gains on low-resource languages, even if it means a tiny dip in performance on the high-resource ones.
Data work has to move past simply hoovering the web. Real partnerships with community organizations can produce culturally rich, ethically sourced collections. For languages with deep oral traditions, weaving in speech data and building solid spoken-to-written bridges can route around the digital text bottleneck entirely. These efforts are slow and hands-on, but they honor a simple principle: the tech ought to bend toward human needs, not the other way around.
My own work has drifted toward cataloging failure modes—not to patch them all at once, but to map the labyrinth’s shape. Each error is a signpost: a morphological collapse in Turkish, a register mismatch in Amharic, a cultural misreading in Japanese. Together they sketch a picture of a world where our tools reflect old inequities right back at us. The real question isn’t whether models perform unevenly. It’s what we decide to do once we see the pattern clearly.
Frequently Asked Questions
Why do some languages need so much more data just to reach acceptable performance?
Languages with tangled morphology, flexible word order, or deep cultural embedding demand more examples before a model can lock onto the underlying patterns. A heavily agglutinative language like Finnish generates an explosion of possible word forms compared to English, so the model needs proportionally more data to encounter them and generalize. On top of that, if the available data is narrow in register, the model never picks up the full expressive range of the language.
Can a model truly understand a concept that exists only in one culture?
“Understand” is a heavy word. A model can learn statistical associations around a culture-specific term, but it lacks the lived texture and social context that fill the concept with meaning. It can notice that “ubuntu” tends to appear near words about community and humanity, but it can’t feel the ethical pull the term carries in many Southern African cultures. That gap between surface association and deep meaning is a stubborn source of unevenness.
Is the problem mainly about computing power or about data availability?
Both matter, but data availability is the deeper, knottier issue. More computing power can help a model train more efficiently on what little it has, but if the data is sparse, unrepresentative, or skewed, no amount of computation will conjure missing linguistic knowledge. The digital divide—the fact that so many languages are barely documented online—is a social and historical tangle that technology alone can’t undo.