When a computational system processes text in English, it often shows a fluency and precision that can feel almost uncanny. Hand the same system a sentence in Swahili, or ask it to build a coherent paragraph in Thai, and the results shift—sometimes halting, sometimes riddled with errors that betray a shallow grip on the language’s bones. This unevenness isn’t a random glitch. It’s a predictable outcome rooted in how these systems are built, the data they swallow, and the linguistic features of the languages themselves. Aiko Murakami, a researcher who has spent years studying cross-linguistic computational behavior, approaches this phenomenon not as a failure to fix overnight, but as a window into the deeper puzzle of representing human language in mathematical spaces.

The Data Diet: Why Some Languages Are Better Fed
One of the bluntest explanations for performance gaps sits in the training data. Multilingual systems typically train on enormous corpora scraped from the web, and the web is not linguistically democratic. English, propped up by history and economics, dominates online text. A 2023 study by Common Crawl—which supplies a big chunk of training data for many large-scale systems—found that English accounts for roughly 46% of all documents in their corpus. Russian, German, and French follow, but at single-digit percentages. Languages like Amharic, spoken by over 30 million people, or Quechua, with millions of speakers across the Andes, appear in fractions of a percent—if they appear at all.
This imbalance creates a fundamental asymmetry. During training, the system’s internal parameters adjust overwhelmingly in response to English-language patterns. The model learns to pour its representational capacity into the statistical regularities of English syntax, semantics, and discourse. When it later meets a low-resource language, it has to generalize from this English-heavy foundation, often mapping unfamiliar structures onto a conceptual frame that was never optimized for them. The result is a kind of linguistic myopia: the system sees the new language through the lens of the dominant one.
But the problem isn’t just about volume. The quality of data matters deeply. For high-resource languages, training sets can be pulled from edited publications, parliamentary records, and carefully transcribed speech. For many others, the available text is often noisy—scraped from social media with inconsistent spelling, mixed with other languages in code-switched posts, or simply too scarce to capture the full morphological and syntactic range of the language. A model trained on such data may learn a brittle, surface-level approximation that collapses when it faces a slightly different dialect or register.
Tokenization: The Hidden Gatekeeper
Before any text reaches the core of a multilingual system, it passes through a tokenizer—a component that chops raw text into smaller units, or tokens, that the model can process. The design of this tokenizer has an outsized influence on performance, yet it’s rarely discussed outside specialist circles. Most modern tokenizers use subword algorithms like Byte-Pair Encoding (BPE), which iteratively merge the most frequent character pairs in a training corpus to build a vocabulary of subword units.
Here lies a subtle but powerful bias. The tokenizer’s vocabulary is learned from a specific corpus, and if that corpus is dominated by English, the resulting subword units will be optimized for English morphology. English words like “running” might be split into “run” and “ning,” a decomposition that reflects English’s relatively simple inflectional system. When the same tokenizer encounters a word from a morphologically rich language like Turkish—say, “evlerinizden” (meaning “from your houses”)—it may be forced to break it into many more fragments: “ev,” “ler,” “in,” “iz,” “den.” Each fragment carries less semantic weight, and the model must piece together meaning across a longer sequence of tokens. This increases computational burden and introduces more points of potential error.
Worse, for languages that use non-Latin scripts, tokenization can be catastrophically inefficient. A single Devanagari syllable, which encodes a consonant-vowel combination in one Unicode block, might be split into multiple tokens because the tokenizer was trained primarily on Latin-script data. The model then struggles to learn meaningful representations for these fragments, leading to degraded performance on tasks like translation or sentiment analysis in Hindi or Marathi.

The Tyranny of the Resource-Rich: Cross-Lingual Transfer and Its Limits
One of the central promises of multilingual systems is cross-lingual transfer—the idea that knowledge learned from a high-resource language can lift performance on a low-resource one. This works because the model’s internal representations are shared across languages; when it learns that “cat” and “feline” are related in English, it might also learn that “gato” and “felino” are related in Spanish, even without explicit Spanish training data for that relationship. This transfer is a remarkable feat of abstraction, but it is far from uniform.
The effectiveness of transfer depends heavily on linguistic similarity. A model trained primarily on English and French will transfer knowledge to Spanish relatively well, because these languages share Indo-European roots, similar word order, and overlapping vocabulary. Transfer to Japanese, however, is much harder. Japanese has a fundamentally different syntactic structure (subject-object-verb versus English’s subject-verb-object), a complex system of honorifics that has no direct parallel, and a writing system that mixes three scripts. The model’s English-centric representations simply do not provide a good prior for these phenomena.
Additionally, the very mechanism of transfer can introduce distortions. The model may overgeneralize English-centric patterns, forcing a subject-verb-object template onto a language that uses a different order, or misinterpreting topic-comment structures as subject-predicate. This is not a failure of learning per se, but a consequence of the model’s inductive bias—its built-in tendency to prefer certain types of solutions over others. When the training data is overwhelmingly in one language, that language’s structural properties become the default, and other languages are treated as deviations to be accommodated rather than systems in their own right.
The Curse of Multilinguality
There is a paradoxical phenomenon known as the “curse of multilinguality,” where adding more languages to a model can actually decrease performance on some of them, particularly low-resource ones. This happens because the model’s fixed capacity must now be shared across more languages. The representational space becomes crowded, and languages with less data or less “signal” during training get pushed to the periphery. Their representations become entangled with those of other languages, leading to interference. For example, a model trained on both Hindi and Urdu—closely related languages that share much vocabulary but differ in script and some grammatical structures—might confuse the two, generating Hindi words in an Urdu context or vice versa.
This interference is not merely an inconvenience; it reveals a fundamental tension in multilingual modeling. The very sharing of representations that enables cross-lingual transfer also creates the conditions for negative interference. Balancing these forces requires careful design choices, such as increasing model capacity, adjusting sampling ratios to overrepresent low-resource languages, or introducing language-specific parameters. But these solutions are partial, and the curse remains a stubborn obstacle.

Linguistic Distance and Structural Mismatch
Beyond data and architecture, the intrinsic properties of languages themselves create uneven performance. Languages differ along dozens of axes: morphological complexity, syntactic flexibility, word order rigidity, use of tone, presence of grammatical gender, and so on. A system that excels at English—a language with relatively simple inflectional morphology and strict word order—may stumble when faced with a polysynthetic language like Inuktitut, where a single word can express what requires an entire sentence in English. The model’s internal machinery, tuned to expect a certain density of meaning per token, is suddenly confronted with tokens that carry vastly more grammatical information.
Consider the challenge of modeling free word order. In English, the position of a noun relative to a verb is a strong cue for determining subject versus object. In languages like Russian or Latin, case marking on the nouns themselves carries this information, allowing words to be rearranged for emphasis or style without changing grammatical roles. A system that has learned to rely heavily on positional cues will be baffled by a language where position is flexible and morphological cues are essential. It must learn to attend to entirely different signals, and this relearning is difficult when the majority of training data reinforces the original strategy.
Script also plays a role that goes beyond tokenization. Languages written in non-Latin scripts often have fewer digital resources overall, compounding the data scarcity problem. But even when data is plentiful, the visual and structural properties of the script can affect learning. Scripts with a high degree of homography—where different characters look similar—can confuse optical character recognition systems that feed into text corpora, introducing noise. Scripts without explicit word boundaries, like Thai, require upstream segmentation that may introduce errors before the model even sees the text.
Evaluation Gaps: Measuring What We Can, Not What We Should
The unevenness we observe is also a product of how we measure performance. Benchmarks for multilingual systems are overwhelmingly designed by English-speaking researchers and reflect English-centric assumptions about what constitutes “good” language understanding. A typical benchmark might test sentiment analysis on product reviews, but the expression of sentiment varies culturally. In some languages, indirectness and hedging are the norm; a direct statement of opinion would be socially aberrant. A system trained to associate certain keywords with positive or negative polarity will misread these culturally conditioned expressions.
Additionally, the tasks themselves may be mismatched to the language. Question answering datasets are often created by translating English questions into other languages, which imports English syntactic structures and cultural assumptions. A question about “breaking a leg” in a theatrical context, translated literally into a language where the idiom does not exist, tests the model’s ability to handle nonsensical input rather than its genuine understanding of the target language. The performance gap we measure is, in part, a gap between the language’s natural use and the artificial tasks we impose on it.
There is also a troubling lack of benchmarks for many languages altogether. While English has hundreds of diverse evaluation sets, many languages have none. Researchers often rely on zero-shot cross-lingual transfer evaluation, where a model trained on English task data is tested on another language without any fine-tuning. This setup inherently favors languages that are typologically and culturally close to English, exaggerating the performance gap.
Morphological Richness and the Sparsity Problem
Morphology—the way words change form to express grammatical relationships—is a major axis of variation. English has relatively simple morphology: a few irregular plurals, three verb forms (walk, walks, walked), and some derivational affixes. In contrast, languages like Turkish, Finnish, or Swahili have extensive agglutinative morphology, where long strings of suffixes attach to a root, each adding a specific grammatical meaning. A single Turkish verb root can generate thousands of distinct forms.
For a system that learns by observing co-occurrence patterns, this morphological richness creates a data sparsity problem. Each inflected form appears rarely, making it hard for the model to learn that “evimde” (in my house) and “evinde” (in your house) share a common root and differ only in the possessive suffix. The model may treat them as entirely separate words, failing to generalize across the paradigm. This leads to poor performance on tasks like machine translation, where producing the correct inflected form is essential, or text generation, where the output may contain morphologically ill-formed words.
Agglutinative languages also challenge the assumption, baked into many architectures, that a word is a coherent unit of meaning. When a single word can express what English requires a phrase to convey, the alignment between source and target languages becomes much harder to learn. The model must learn to map a single token in one language to multiple tokens in another, a process that is inherently more complex than one-to-one or one-to-few mappings.
Script and Orthography: The Surface Barrier
Script differences introduce another layer of complexity. Many multilingual systems use a shared vocabulary of subword tokens that must accommodate multiple scripts. This forces the tokenizer to allocate vocabulary slots to characters and character sequences from many writing systems, reducing the granularity available for any single script. Languages with larger character inventories—like Chinese, with thousands of distinct characters—suffer disproportionately, as their characters must be represented by multiple tokens, fragmenting meaning.
Orthographic depth also matters. English has a deep orthography, where the relationship between spelling and pronunciation is complex and often irregular. Other languages, like Spanish or Finnish, have shallow orthographies with more consistent mappings. A system trained primarily on English may learn to rely on memorizing whole-word spellings rather than decoding phonological patterns. When faced with a shallow orthography, it may fail to exploit the regularities that native speakers use effortlessly, leading to poorer performance on tasks that require phonological awareness, such as rhyming or pun detection.
Cultural and Contextual Knowledge
Language is not just a system of rules and vocabulary; it is a carrier of cultural knowledge. Understanding a text often requires knowing about the world it refers to—historical events, social norms, popular figures, and shared metaphors. Multilingual systems absorb this knowledge from their training data, which is overwhelmingly produced by and for Western, educated, industrialized, rich, and democratic (WEIRD) societies. When processing a text from a different cultural context, the system lacks the necessary background knowledge.
For instance, a system might perform well on English-language questions about Thanksgiving or the Magna Carta but fail on questions about Diwali or the Ramayana in Hindi, not because it cannot process the language, but because it lacks the cultural knowledge to interpret the references. This is a form of unevenness that is not strictly linguistic but is deeply intertwined with language. The model’s knowledge base is culturally skewed, and this skew maps onto language performance because language and culture are inseparable.
Frequently Asked Questions
Why do some languages with many speakers still perform poorly?
Performance depends less on the number of speakers and more on the volume and quality of digital text available for training. Languages like Bengali or Punjabi have hundreds of millions of speakers but relatively limited web presence compared to English or German. Additionally, if the available text is noisy—from social media, for example—or mixed with other languages, the model may struggle to learn clean representations. Speaker population does not guarantee a large, high-quality digital corpus.
Can simply adding more data for a low-resource language fix the problem?
Adding more data helps, but it is not a complete solution. The model’s architecture and training procedure also matter. If the tokenizer is poorly suited to the language’s script or morphology, more data will not fully compensate. Additionally, if the model’s capacity is fixed, adding data for one language may crowd out representations for others, exacerbating the curse of multilinguality. A combined approach—better data, language-appropriate tokenization, and capacity management—is needed.
Why do some languages with similar structures to English still underperform?
Even typologically similar languages can suffer if their digital presence is small or if the training data is dominated by English to the point that the model overfits to English-specific patterns. For example, Dutch and English share many structural features, but if Dutch data is a tiny fraction of the training set, the model may treat Dutch as a minor variant of English rather than learning its distinct properties. Subtle differences in word order, idiomatic expressions, or preposition usage can then cause errors.
Is the performance gap narrowing over time?
In some respects, yes. Researchers are actively developing techniques like targeted data augmentation, better tokenization strategies, and language-specific fine-tuning. However, the gap is not closing uniformly. High-resource languages continue to improve rapidly, while many low-resource languages see only marginal gains. The underlying structural challenges—linguistic distance, script differences, cultural knowledge gaps—require more than just scaling up existing methods; they demand fundamentally new approaches to multilingual representation.
The uneven performance of multilingual systems is not a temporary glitch on the path to universal language mastery. It is a reflection of deep asymmetries in the digital world, in the design of our computational tools, and in the very nature of human languages. Understanding these asymmetries requires a scholarly patience, a willingness to look beyond aggregate metrics and examine the specific ways each language is represented—or misrepresented—inside the model. Only then can we begin to build systems that treat all languages not as variations on a dominant theme, but as complete and sovereign linguistic worlds.