Every language model carries a linguistic fingerprint from its pre-training corpus. Not a design choice—a statistical residue. Patterns of word order, morphological complexity, and agreement seep in without anyone explicitly programming them. For those of us who study morphosyntax across languages, the real question isn’t whether pre-training data matters. It’s how tightly it constrains the generalizations a model can actually form. Train a model mostly on English news text, and it develops one kind of inductive bias. Train it on a balanced sample of fifty languages, and you get something else entirely. Test both on subject-verb agreement or case marking, and the results mirror the distributional quirks of the training data far more than any abstract universal grammar.
This article examines that relationship—between what a model is fed and how it handles morphosyntax. We’ll look at concrete examples from dependency parsing, agreement attraction, and cross-linguistic transfer, drawing on recent evaluation benchmarks and controlled experiments. The goal is a clear, evidence-backed picture of how data provenance shapes what a model can and cannot learn about sentence structure.
What Pre-Training Data Encodes
Pre-training data isn’t a neutral snapshot of a language. It’s a curated, often web-scraped collection that leans heavily on certain registers, dialects, and syntactic constructions. For English, that means a glut of declarative sentences, news-style prose, and rigid subject-verb-object order. For morphologically richer languages, the picture gets patchier. A 2022 audit of multilingual corpora found that even in so-called “balanced” datasets, average sentence length and morphological density varied by a factor of two across languages (Kreutzer et al., 2022). These imbalances harden into the model’s default expectations.
Take subject-verb agreement. In English, it’s sparse—only the third-person singular present tense gets an overt suffix. A model trained mostly on English learns to treat agreement as a low-weight cue. Fine-tune that same architecture on a language with lush agreement morphology, like Swahili or Basque, and the pre-trained weights keep tugging it back toward the English-like pattern. The model has to fight its own initial bias. How well it succeeds depends on how much target-language data you supply during fine-tuning.

Morphological Richness and the Data Hunger Problem
Languages encode grammatical information in wildly different ways. Isolating languages like Mandarin lean on word order and particles. Agglutinative languages like Turkish stack suffix after suffix onto a single stem. Train a model on a corpus dominated by isolating languages, and its internal representations will be poor at segmenting and interpreting complex morphemes. That’s not an architecture failure. It’s a direct consequence of the pre-training distribution.
Empirical work backs this up, and it shows that the “data hunger” of morphologically rich languages isn’t uniform. A 2021 study on dependency parsing across 30 languages found that the amount of pre-training data needed to hit a given performance threshold correlated strongly with morphological complexity (Üstün et al., 2021). Languages with more inflectional categories needed more tokens to reach the same attachment scores. The takeaway: pre-training data quantity and diversity aren’t just about scale. They’re about covering the morphological paradigms a model needs to internalize.
Case Study: Agreement Attraction in Slavic Languages
Agreement attraction happens when a verb mistakenly agrees with a nearby noun instead of its actual subject. In English, it’s relatively rare and well-studied. In Slavic languages, where agreement morphology is richer and word order flexes more freely, the patterns get messier. A model pre-trained on a corpus with thin Slavic representation will show higher attraction rates and more erratic behavior when you test it on sentences with non-canonical word orders.
In controlled experiments, researchers have tweaked the proportion of Slavic-language data in the pre-training mix and measured the resulting error rates on number-agreement tasks. The findings are consistent: error rates drop sharply as the target language’s representation increases, but the benefit plateaus once that language makes up roughly 5–8% of the total corpus. Beyond that, more data brings diminishing returns for agreement accuracy. It suggests that a relatively modest amount of well-curated data can fix major syntactic blind spots—provided that data covers the relevant morphological paradigms.
Cross-Linguistic Transfer: When Pre-Training Helps and Hurts
Cross-linguistic transfer often gets presented as a clear win for multilingual pre-training: train on many languages, and the model can use shared structure to boost performance on low-resource ones. The reality is messier. Transfer helps when the source and target languages share typological features that are well-represented in the pre-training data. When they don’t, pre-training can actively get in the way.
A clean example comes from case marking. German and Russian use case to signal grammatical relations. English relies almost entirely on word order. A model pre-trained on an English-heavy corpus develops a strong word-order bias. Fine-tune it on German, and it may ignore case cues in favor of linear order, leading to systematic errors on sentences where the object comes before the subject. This interference effect has been documented in several probing studies. The model’s attention heads keep prioritizing the first noun phrase as the agent, even when case marking says otherwise.

Quantifying Interference with Typological Distance
One way to predict cross-linguistic interference is to measure the typological distance between the pre-training language distribution and the target language. The World Atlas of Language Structures (WALS) gives you feature vectors for hundreds of languages—word order, morphological type, alignment. Compute the cosine distance between the average feature vector of the pre-training corpus and the feature vector of the target language, and you get an estimate of the mismatch a model will face.
In practice, this means a model pre-trained on a corpus dominated by SVO, isolating languages will struggle most with SOV, agglutinative languages. The interference isn’t just about vocabulary overlap. It’s structural. Fine-tuning can soften the problem, but it can’t fully overwrite the pre-trained biases, especially when the fine-tuning dataset is small. This has direct implications for evaluation methodology: if a benchmark doesn’t control for typological distance from the pre-training data, it may overestimate a model’s cross-linguistic generalization ability.
Evaluation Benchmarks and Their Blind Spots
Standard evaluation benchmarks for morphosyntax, like BLiMP and SyntaxGym, lean heavily toward English. Translate or adapt them to other languages, and they often miss the phenomena that are most challenging for models trained on English-dominant data. A benchmark that tests subject-verb agreement in English won’t probe for ergative case alignment or noun class agreement—both critical for many non-Indo-European languages.
Recent efforts to build typologically-informed benchmarks have uncovered systematic gaps. The SIGMORPHON shared tasks on morphological inflection and the UniMorph project provide cross-linguistic coverage, but they focus on form generation rather than syntactic evaluation. A more integrated approach would combine morphological and syntactic probes within a single framework, letting researchers measure how pre-training data composition affects the full pipeline from morpheme recognition to sentence interpretation.
Data Mixtures and the Long Tail of Languages
Even within a single language, pre-training data isn’t homogeneous. Web text overrepresents formal registers, standard dialects, and certain demographic groups. For morphosyntax, that means non-standard agreement patterns, dialectal variation, and colloquial constructions get short shrift. A model trained on such data may ace edited newswire but stumble on social media text or regional varieties. This is a data problem, not an architecture problem, and it demands careful documentation of what the pre-training corpus actually contains.
One practical step is to release data sheets that describe the linguistic composition of pre-training corpora. These sheets should include information on language distribution, register, and morphological density. Without that documentation, it’s impossible to tell whether a model’s poor performance on a given construction reflects a genuine difficulty or just a mismatch between the evaluation data and the pre-training distribution.

Practical Implications for Morphosyntax Research
For researchers building evaluation suites or probing classifiers, pre-training data is a variable you have to control. When you compare two models on a morphosyntactic task, performance differences may reflect pre-training data differences rather than architectural innovations. Reporting the language distribution and corpus composition is just as important as reporting hyperparameters.
One concrete recommendation: stratify evaluation results by typological feature. Instead of a single accuracy score for “subject-verb agreement,” report separate scores for head-initial and head-final languages, for languages with and without case marking, and for languages with different morphological complexity indices. That stratification makes it possible to see whether a model’s errors cluster in typological regions that were underrepresented in its pre-training data.
Fine-Tuning as a Partial Remedy
Fine-tuning on target-language data can partially correct pre-training biases, but it’s not a full fix. The pre-training distribution keeps exerting influence, especially in low-data regimes. One study on dependency parsing across 50 languages found that even after fine-tuning on thousands of target-language sentences, the model’s attachment preferences were still shaped by the pre-training language distribution. This “pre-training shadow” was most pronounced for rare constructions and long-distance dependencies.
For practitioners, that means fine-tuning data should be curated to overrepresent the constructions known to cause trouble. If a language has flexible word order, the fine-tuning set should include examples of all major orders, not just the most frequent one. If a language has complex agreement morphology, the fine-tuning set should include examples with non-local agreement controllers. These targeted interventions can shrink the pre-training shadow, though they can’t erase it entirely.
FAQ
Why does pre-training data composition matter more for morphosyntax than for other linguistic levels?
Morphosyntax involves systematic dependencies between words and morphemes that are often language-specific. Unlike lexical semantics, where cross-linguistic overlap can be substantial, morphosyntactic patterns vary widely. A model pre-trained on a corpus that lacks a particular pattern—such as ergative case marking or polypersonal agreement—has no opportunity to learn the underlying generalizations. The result is not just lower accuracy, but qualitatively different error types that reflect the model’s native-language bias.
Can a model learn morphosyntax that is absent from its pre-training data?
In principle, a model can acquire new morphosyntactic patterns through fine-tuning, but the learning is often shallow. Studies on low-resource morphological inflection show that models can memorize frequent forms but struggle to generalize to unseen stems or novel combinations of morphemes. True productivity—the ability to inflect a novel verb correctly—requires exposure to a sufficient number of paradigm cells during pre-training. Without that exposure, the model treats morphology as a lexical property rather than a combinatorial system.
How can researchers evaluate whether pre-training data is the cause of a model’s errors?
The most direct method is to compare models pre-trained on different data distributions while holding architecture constant. If a model trained on a corpus with more target-language data shows better morphosyntactic generalization, the pre-training data is likely the causal factor. Another approach is to use probing classifiers to measure the extent to which morphosyntactic features are encoded in the model’s representations at different layers. If a feature is absent or weakly encoded, and the pre-training data contains few examples of that feature, the data is the probable bottleneck.
What can be done to improve morphosyntactic generalization across languages?
Improving generalization requires deliberate data curation. Pre-training corpora should be sampled to ensure coverage of diverse morphological systems, not just diverse languages. For evaluation, benchmarks should include typologically stratified test suites that isolate specific morphosyntactic phenomena. Finally, reporting standards should require documentation of pre-training data composition, so that the field can distinguish between architectural limitations and data-induced biases.
Next Steps for the Blog
This article has focused on how pre-training data composition affects morphosyntactic behavior. A natural follow-up would examine how fine-tuning data interacts with pre-training biases: does fine-tuning on a typologically diverse set of languages reduce the pre-training shadow, or does it simply overlay new biases on top of the old ones? Another direction is to investigate whether data augmentation techniques, such as back-translation or morphological reinflection, can compensate for gaps in the pre-training distribution. These questions will shape the next phase of empirical work on this blog.