Introduction: The Data Diet of a Language Model
When a model handles English syntax with ease but stumbles on a simple German relative clause, the problem usually isn’t the architecture. It’s the training data. The mix of languages, the types of texts, and the presence or absence of specific grammatical patterns in the pre-training corpus all leave an indelible mark on what the model can and cannot do. This article examines how the composition of pre-training data determines model behavior, with a focus on morphosyntax and the challenges of cross-linguistic generalization. We’ll look at what happens when one language dominates the data, why that creates systematic blind spots, and how evaluation methods need to change to reveal them.

The Primacy of the Training Distribution
Models don’t learn grammar from explicit rules. They absorb statistical regularities from enormous text corpora. The resulting behavior is a mirror of the pre-training data distribution. If a corpus is 90% English, the model will develop strong representations for English subject-verb agreement, article usage, and word order. For a language like Hindi, which might make up a fraction of a percent of the same corpus, the model’s morphosyntactic abilities will be correspondingly weak.
This isn’t a question of model size. Even massive models with billions of parameters show a tight link between data frequency and performance. The model learns to allocate its representational space to patterns that maximize predictive accuracy during training. Rare patterns get sidelined as noise. It’s a sensible strategy for minimizing loss, but a poor one for building a general-purpose linguistic system.
Data Imbalance and Morphosyntactic Blindness
Morphosyntax—the intersection of word formation and sentence structure—is especially vulnerable to data imbalance. Consider case marking. English has a vestigial case system, limited mostly to pronouns (“I” vs. “me”). Russian or Finnish, by contrast, use case extensively, and it interacts with number, gender, and animacy. A model pre-trained on overwhelmingly English data has seen thousands of examples of English pronoun case but may have never encountered the Finnish partitive case. When tested on Finnish, it doesn’t just perform poorly; it often imposes English-like syntactic strategies, treating Finnish words as if they followed English word order constraints.
This is a direct consequence of the training distribution. The model learned that word order is a reliable cue for grammatical function because, in its primary training language, it is. Without sufficient exposure to case-rich languages, the model never learns to rely on morphological cues. The result is a systematic, data-induced blind spot.
Cross-Linguistic Generalization: A Mirage?
Researchers often test whether models can transfer knowledge from high-resource languages to low-resource ones—so-called cross-linguistic generalization. The results are uneven. For tasks like named entity recognition, shared subword tokens or structural similarities sometimes provide a bridge. For morphosyntax, the outlook is less encouraging.
A model trained on English and French might show some ability to parse Spanish, another Romance language. But test the same model on Turkish, with its agglutinative morphology and vowel harmony, and performance often drops to chance. The pre-training data simply didn’t supply the necessary building blocks. The model has no internal representation of vowel harmony because it never had a reason to learn one.

The Role of Subword Tokenization
Subword tokenization algorithms like Byte-Pair Encoding (BPE) are often presented as a solution for data sparsity. By breaking words into smaller units, models can share representations across languages. This works reasonably well for languages with similar scripts and morphological structures. A model can learn that “-ing” in English and “-ung” in German both signal nominalization, provided both appear in the training data.
But tokenization itself is a product of the pre-training data. The BPE vocabulary is built from the training corpus. If a language is underrepresented, its subword units get fragmented into smaller, less meaningful pieces. A Finnish word like “talossanikin” (“in my house, too”) might be split into a dozen subword tokens, while an English equivalent remains a single token. This fragmentation makes it harder for the model to learn morphosyntactic patterns. The model sees a sequence of seemingly unrelated tokens rather than a coherent morphological structure.
Evaluation Pitfalls: Measuring What the Data Dictates
Standard evaluation benchmarks often conceal the role of pre-training data. A model can score well on a multilingual benchmark simply because the benchmark’s language distribution mirrors the training data. When researchers construct adversarial test sets that probe specific morphosyntactic phenomena—like subject-verb agreement across long-distance dependencies in low-resource languages—performance drops sharply.
This exposes a fundamental tension. Evaluation sets drawn from the same distribution as the training data measure memorization and pattern matching, not linguistic competence. To understand what a model has actually learned, we need evaluation data that systematically varies linguistic features and controls for corpus frequency. Only then can we separate the model’s inherent capabilities from the artifacts of its data diet.
Case Study: Agreement in Slavic Languages
Slavic languages provide a useful testbed. They have complex agreement systems involving gender, number, and case. Some, like Russian, have three genders; others, like Slovenian, also have a dual number. A model pre-trained on mostly English and Western European data will have limited exposure to these categories. When asked to predict the correct form of a Russian adjective modifying a feminine noun in the instrumental case, the model often defaults to the nominative masculine form—the most frequent form in its training data.
This isn’t a random error. It’s a systematic bias induced by the training distribution. The model learned that the most frequent form is the safest bet. Without enough evidence to learn the agreement rules, it falls back on a frequency heuristic. This behavior holds across model sizes and architectures, which points to a data problem, not a capacity problem.
Data Documentation and Model Transparency
Understanding model behavior means understanding the training data. Yet for many models, the pre-training data composition remains opaque. We know the total token count, but not the language breakdown, the genre distribution, or the presence of specific morphosyntactic constructions. This opacity makes it hard to predict where a model will fail.
Efforts to document training data, such as the BigScience Workshop’s work on the ROOTS corpus, mark a step forward. Detailed metadata on language representation, data sources, and processing steps let researchers better anticipate model weaknesses. Without such documentation, every failure is a surprise, and every success is hard to replicate.

Practical Implications for Multilingual Systems
For anyone building applications on top of language models, the data-behavior link has real consequences. A model that handles English customer service queries well may fail badly on code-switched input or morphologically rich languages. The answer isn’t always a bigger model. Often, it’s a more carefully curated pre-training corpus.
Targeted data augmentation can help. Adding even a small amount of high-quality, morphologically diverse text can improve performance on specific phenomena. But it’s not a cure-all. The model’s internal representations are shaped by the entire data distribution. A few thousand Finnish sentences won’t override the statistical dominance of English. The model may learn to recognize some Finnish patterns, but it won’t develop a reliable Finnish grammar.
Evaluation Beyond Accuracy
To really understand the impact of pre-training data, we need evaluation methods that go beyond aggregate accuracy. Probing classifiers, minimal pair tests, and targeted syntactic evaluations can reveal what linguistic features a model has actually learned. For example, a probing classifier trained on model representations can test whether the model encodes grammatical gender. If the classifier performs well on languages with high representation but fails on low-resource languages, we have direct evidence of data-induced blind spots.
These methods also let us track how representations change as we vary the pre-training data. By training multiple models on different data mixtures, we can map the relationship between data frequency and linguistic competence. This is slow, expensive work, but it’s necessary for building models that serve more than a handful of high-resource languages.
FAQ
Why does a model make more errors in morphologically rich languages?
The main reason is data imbalance. Pre-training corpora are dominated by English and a few other high-resource languages. Morphologically rich languages like Turkish, Finnish, or Arabic appear far less often. The model therefore has fewer chances to learn the complex inflectional patterns and agreement rules that characterize these languages. It often defaults to the most frequent pattern from its training data, which is usually an English-like structure.
Can we fix data imbalance by simply adding more data from low-resource languages?
Adding more data helps, but it’s not a complete fix. The model’s learning dynamics are influenced by the overall data distribution. If English still dominates, the model may treat the new data as noise or fail to integrate it into a coherent grammatical system. The quality of the added data also matters. Noisy, poorly curated text can introduce new errors. A better approach combines targeted data augmentation with architectural modifications that encourage the model to share representations across languages more effectively.
How can I tell if a model’s errors are due to pre-training data or some other factor?
Look for systematicity. If errors cluster in specific languages or linguistic phenomena that are known to be underrepresented in the training data, data imbalance is a likely cause. Compare performance across languages with different levels of representation in the training corpus. If performance correlates with data frequency, the data is a primary driver. Controlled experiments, such as training models on different data mixtures and comparing their behavior, provide the strongest evidence.
Conclusion: The Data We Choose, The Models We Get
Pre-training data isn’t just fuel for a model; it’s the blueprint. Every morphosyntactic pattern a model can recognize, every cross-linguistic generalization it can make, and every error it systematically produces is a direct consequence of the texts it was trained on. As the field moves toward larger and more opaque training sets, the need for careful data documentation and targeted evaluation grows. The goal isn’t to build a model that performs well on a single leaderboard, but to understand what linguistic knowledge has been acquired and what has been left out. That understanding begins and ends with the data.