Why Multilingual Models Perform Unevenly Across Languages

Ask the same question in English and then in Swahili, and you might get two very different answers. The English response arrives crisp, confident, and well-structured. The Swahili one? Sometimes it’s thinner, a little off-target, or phrased in a way that no native speaker would choose. This lopsidedness isn’t a random bug. It grows out of a tangle of data imbalances, linguistic distance, evaluation blind spots, and the way shared model architectures quietly favor the languages they’ve seen most. To see why some languages get the short end of the stick, we need to look at what goes into training, how meaning gets encoded, and what we actually measure when we call one language “better supported” than another.

Globe with interconnected digital nodes representing global language networks

The Data Diet: What You Feed the Model Shapes Everything

Multilingual models eat huge, scraped collections of web pages, digitized books, news articles, and whatever else is publicly available. The mix of languages in that diet is the single biggest predictor of how well the model will perform later. English, Chinese, Spanish, German—these languages flood the training pot with billions of tokens. Meanwhile, many languages spoken by millions of people barely show up. One audit of a popular multilingual corpus found English alone made up about 46% of all tokens, while more than 90% of the languages contributed less than 1% each. For those low-resource languages, the model simply hasn’t seen enough natural syntax, everyday idioms, or specialized vocabulary to learn them reliably.

Quality matters as much as quantity. High-resource languages often come with professionally edited text—news reports, academic papers, technical manuals—that’s consistent and factually tight. Low-resource language data, on the other hand, is often scraped from social media, informal forums, or machine-translated pages. You get non-standard spelling, code-switching, sentence fragments, and translation artifacts that don’t match how people actually speak and write. Feed a model noisy data, and it learns noisy patterns. The output ends up grammatically awkward, semantically shallow, or stylistically wrong for the context.

The Long Tail of Digital Scarcity

Linguists count over 7,000 living languages, but fewer than 100 have a substantial digital footprint. This “digital language divide” isn’t just about speaker numbers. Economics, internet access, and colonial histories all shape which languages pile up written records online. Javanese has more than 80 million speakers but remains severely underrepresented in training corpora—much daily communication happens orally or in scripts less common in digital spaces. Meanwhile, Finnish or Catalan, with smaller speaker bases but strong institutional backing, enjoy disproportionately rich textual resources. The model’s uneven performance mirrors this uneven digital landscape.

Books stacked unevenly, symbolizing imbalanced language resources

Linguistic Distance and the Shared Representation Tangle

Most multilingual models cram all languages into a single shared vector space. The idea is elegant: knowledge learned in one language can transfer to others—a trick called cross-lingual transfer. The assumption is that concepts, relationships, even syntactic functions sit in roughly the same region of the embedding space no matter which surface language you use. In practice, how well transfer works depends heavily on structural similarity. Languages that share typological features—word order, morphological complexity, writing system—cluster more tightly, making transfer smoother. Typologically distant languages, especially those far from the high-resource anchors, can end up poorly aligned, their representations scattered or squeezed into suboptimal corners of the space.

Morphology throws a wrench in the works. English leans on word order and auxiliary verbs to express grammatical relationships. Turkish, by contrast, packs meaning into long strings of suffixes: a single word can carry what English needs a whole phrase for. When a model trained mostly on English-like patterns meets a morphologically rich language, the tokenization strategy—usually subword units optimized for English—can chop roots and affixes into pieces that obscure meaning. The model may not recognize that three different surface forms are inflections of the same lemma, or that a suffix carries tense and aspect information. This fragmentation leads to sparser, less coherent representations and weaker performance on translation, summarization, and question answering.

Script and Tokenization Mismatches

Writing systems add another layer of unevenness. Languages using Latin scripts benefit from tokenizers heavily optimized on English and other European languages. Scripts like Devanagari, Arabic, or Ge’ez often need more tokens per word because the tokenizer’s vocabulary was built with different character distributions in mind. This “token fertility” problem means the same semantic content eats up more of the model’s fixed context window, leaving less room for subtle reasoning. It also raises computational cost per query, which can lead to truncated or less detailed responses even when the underlying knowledge is there.

Evaluation Gaps: What We Measure and What We Miss

Performance claims about multilingual models are only as solid as the benchmarks used to test them. The most cited evaluation suites—XNLI for natural language inference, XQuAD for question answering, FLORES for translation—cover a limited set of languages, often fewer than 15. Even when a benchmark claims 100 languages, the test sets are frequently built by translating English source material, sometimes with machine translation and light human post-editing. This injects “translationese”: artifacts of the source language that make the task easier or harder in ways that don’t reflect genuine linguistic competence. A model might ace a translated Swahili benchmark not because it understands Swahili deeply, but because it has learned to map translationese patterns back to English.

Standard benchmarks also measure narrow slices of language ability—factual accuracy, entailment judgment, exact match span extraction. They rarely touch pragmatic fluency, cultural appropriateness, or dialectal variation. A model that scores well on a formal Hindi benchmark may still churn out stilted, overly Sanskritized output that alienates everyday speakers. In languages with strong diglossia, like Arabic, where the written standard differs sharply from spoken vernaculars, benchmark scores can paint an overly rosy picture of real-world usability. The unevenness users feel often lives precisely in these unmeasured dimensions.

Magnifying glass over a multilingual document, highlighting evaluation scrutiny

The Feedback Loop of Resource Allocation

Benchmark results steer research attention and funding. Languages that score poorly attract less follow-up work, fewer curated dataset efforts, and minimal hyperparameter tuning. This creates a self-reinforcing cycle: low resource leads to low performance, which leads to low investment, which perpetuates low resource. Breaking the cycle takes deliberate effort—collecting native-speaker data, designing culturally grounded evaluations, and optimizing model architectures specifically for underrepresented language families, not just scaling up the same English-centric recipe.

Architectural Constraints and Capacity Allocation

A model’s total capacity—its parameter count—is finite. When that capacity is shared across many languages, a quiet competition kicks in. High-resource languages, with their flood of training signal, tend to hog the parameters. The model allocates more of its representational budget to patterns that reduce loss on the bulk of the training data, which overwhelmingly comes from a handful of languages. Low-resource languages get squeezed into the remaining representational corners, often forced to reuse suboptimal patterns borrowed from unrelated high-resource languages.

This effect is sometimes called the “curse of multilinguality.” Adding more languages to a model can actually degrade performance on existing low-resource languages, as the shared capacity gets further diluted. Researchers have noticed that for very low-resource languages, a dedicated bilingual model often outperforms a massive multilingual one, simply because the bilingual model can devote all its parameters to the relevant pair without interference. The trade-off between breadth and depth isn’t easily resolved; it’s baked into the design choice of sharing parameters across languages.

Attention Imbalance During Training

Training schedules matter too. Many multilingual models use a uniform sampling strategy over languages, but the data within each language is still heavily skewed. Even if the model sees an equal number of batches for each language, the batches for low-resource languages may contain repetitive, low-diversity text because the available corpus is small. The model overfits to the few patterns it sees, memorizing phrases rather than learning productive rules. When it faces novel input in that language, it falls back on memorized fragments or defaults to English-like structures, producing output that feels generic or off-target.

Sociolinguistic Factors: Prestige, Standardization, and Code-Mixing

Language isn’t a uniform, static system. It varies by region, social context, and medium. Multilingual models are typically trained on the “standard” variety of each language—the form used in official documents, news media, and textbooks. Yet many speakers communicate primarily in non-standard dialects, mixed codes, or colloquial registers. When a user queries in Moroccan Arabic (Darija) or Nigerian Pidgin, the model may have little to no training data in that specific variety. It tries to map the input to the closest standard language it knows, often Modern Standard Arabic or English, losing the subtleties of the original query and responding in a register that feels foreign to the user.

This mismatch is especially sharp in multilingual societies where code-switching is the norm. A speaker of Hindi and English might naturally blend the two in a single sentence, but the model—trained to separate languages into distinct buckets—struggles to process the mixed input coherently. The result can be a response that ignores one language entirely, misinterprets the switch as an error, or generates a jarring mix of registers. The model’s uneven performance across languages, then, isn’t only about individual languages in isolation; it’s about the real-world ways people fluidly combine them.

Historical and Political Dimensions

The uneven digital presence of languages isn’t a neutral fact; it reflects historical patterns of colonization, economic inequality, and language policy. Many African and Indigenous languages were systematically excluded from education, publishing, and media for generations, leaving a legacy of sparse written resources. When multilingual models inherit these imbalances, they risk amplifying them—providing excellent service in former colonial languages while offering degraded experiences in local languages, potentially accelerating language shift. This isn’t a technical problem alone; it’s a sociotechnical one that requires engagement with linguists, communities, and policymakers.

What Can Be Done: Targeted Interventions

Addressing uneven performance takes a multi-pronged approach that goes beyond simply collecting more data. For many languages, the available text is already scraped; what’s missing is curated, diverse, and representative data. Community-driven data collection projects—documenting oral histories or digitizing local literature—can provide high-quality training material that reflects genuine usage. Just as important is building evaluation benchmarks that measure what matters to speakers: politeness strategies, domain-specific terminology, dialect comprehension, and cultural relevance.

On the technical side, researchers are exploring modular architectures that allocate dedicated capacity to language-specific components while retaining shared cross-lingual knowledge. Adapter layers, sparse expert mixtures, and language-specific tokenizers can help ease the capacity dilution problem. Training strategies that oversample low-resource languages or use contrastive objectives to sharpen cross-lingual alignment have shown promise in narrowing the performance gap. None of these is a silver bullet, but together they point toward more equitable multilingual systems.

The Role of Native Speaker Involvement

Perhaps the most underappreciated factor is the involvement of native speakers in the development cycle. When the teams building and evaluating models are dominated by speakers of high-resource languages, subtle errors in low-resource languages go unnoticed. Native speakers can spot not only grammatical mistakes but also pragmatic failures—when a response is technically correct yet socially inappropriate. Their feedback is essential for creating models that serve communities rather than merely processing their languages.

FAQ

Why do some languages with many speakers still perform poorly?

Speaker population doesn’t directly translate to digital text availability. Languages like Punjabi or Javanese have tens of millions of speakers but limited online written content due to oral traditions, use of non-dominant scripts, or preference for a colonial language in formal writing. The model’s performance depends on the quantity and quality of written training data, not the number of speakers.

Does adding more languages always hurt performance on existing ones?

Not always, but there’s a known trade-off. When model capacity is fixed, adding languages can dilute the representational budget, especially for low-resource languages that already struggle to claim enough parameters. However, adding typologically similar languages can sometimes improve performance through positive transfer. The effect depends on the specific languages, model architecture, and training strategy.

Can a model be equally good at all languages?

In theory, with unlimited capacity and perfectly balanced, high-quality data, a model could approach parity. In practice, the uneven distribution of digital text, the diversity of linguistic structures, and the limitations of current architectures make perfect equality extremely difficult. The goal isn’t absolute parity but reducing the gap to a point where the model is genuinely useful for speakers of all languages, not just a privileged few.