How Pre-Training Data Shapes Morphosyntactic Behavior: An Empirical View
If you want to know why a model keeps tripping over dative alternation in German or ergative alignment in Hindi, don’t start with the attention heads. Start with the pre-training corpus. The composition of that data—which languages show up, in what genres, and how the tokenizer chops them into pieces—is the single biggest lever on morphosyntactic behavior. This article walks through the evidence. We’ll look at subject-verb agreement in Basque, noun-class marking in Swahili, and differential object marking in Turkish, drawing on controlled experiments that separate data effects from architectural choices. For anyone evaluating cross-linguistic generalization, this isn’t background reading. It’s the foundation.

Data Provenance as a Predictor
It’s easy to treat a model as a black box that either “knows” a language or doesn’t. The reality is messier. A model trained on 90% English and 10% everything else will develop a strong nominative-accusative prior. That prior bleeds into every prediction. When the same model then encounters Basque—a language with ergative-absolutive alignment—the prior actively misleads it. The model isn’t incapable of learning ergative patterns. It just never saw enough of them to override the dominant signal.
Researchers at the Center for Language and Cognition Groningen tested this directly. They varied the proportion of morphologically rich languages in the pre-training mix and measured agreement accuracy. Bumping a low-resource language from 0.1% to 2% of the corpus produced a 23-percentage-point gain. Not a subtle shift. And the gains tapered off after that, which tells you something important: there’s a threshold, and it’s language-specific. More data helps until it doesn’t.
Tokenization: The Hidden Variable
Pre-training data doesn’t just determine what a model sees. It determines how the model sees it. Tokenization algorithms—BPE, SentencePiece, Unigram—are trained on the same corpus, and their subword splits directly shape morphological learning. In agglutinative languages like Turkish or Finnish, a tokenizer that splits morphemes into separate tokens forces the model to learn composition from context. A tokenizer that keeps complex word forms intact may encourage memorization at the expense of generalization.
Differential object marking in Turkish makes this concrete. When a Unigram tokenizer segments accusative-marked objects into stem and suffix tokens, models generalize better to new nouns. Keep those same forms as single tokens, and performance on seen nouns rises—but zero-shot generalization drops by 11 points. The pre-training data’s influence reaches all the way down to the units of analysis.

Genre Bias and Syntactic Generalization
Pre-training data isn’t a uniform soup. It’s a patchwork of genres, registers, and domains. A corpus heavy on newswire and parliamentary transcripts will produce a model that handles declarative, high-register syntax well. Ask it to process informal interrogatives or relative clause attachments, and the cracks show.
A 2024 study in Transactions of the Association for Computational Linguistics put numbers to this. Models trained on genre-balanced corpora outperformed web-crawl baselines on syntactic generalization benchmarks by 8–14 points. The biggest gains came on question formation and relative clause attachment. Total token counts were held constant, so the difference was purely compositional. For evaluation, this means a model’s failure on a syntactic phenomenon may reflect genre underrepresentation, not a hard architectural limit.
Noun-Class Systems and Data Sparsity
Swahili noun classes offer a sharp example. Train a model on 90% English and 10% Swahili, and it will often default to class 9/10 agreement—the most frequent in the Swahili subset—regardless of the actual noun class. This “majority class” effect is a direct consequence of sparsity. The model hasn’t seen enough examples of the less frequent classes to overcome the prior. Increasing Swahili data to 30% reduces the bias but doesn’t eliminate it. Proportional representation alone isn’t enough when intra-language distributions are skewed.
This has practical weight for evaluation. A language can be “present” in the training data while remaining effectively absent for rare constructions. When assessing morphosyntactic capabilities, you need to look past total token counts and examine the distribution of morphological paradigms within that language’s subset.
Cross-Linguistic Transfer: When More Data Helps—and When It Hurts
Cross-linguistic transfer is often sold as a pure win. The evidence is more mixed. For typologically common features—basic SVO order, say—adding more languages to the pre-training mix generally improves accuracy across the board. For features that are typologically rare or parametrically complex, adding languages can introduce interference.
Differential object marking is a useful test case again. Turkish, Hindi, and Spanish all have DOM, but the conditioning factors differ. Turkish uses specificity. Hindi uses animacy and definiteness. Spanish uses animacy and referentiality. A model pre-trained on all three languages shows worse DOM accuracy in Turkish than a model trained on Turkish alone. The conflicting conditioning factors from Hindi and Spanish create a blended representation that fits none of them well. This negative transfer is a direct consequence of pre-training data composition. Fine-tuning alone won’t fix it.
Evaluation Methodology: Controlling for Data Artifacts
If pre-training data shapes model behavior this much, any evaluation that ignores data composition is incomplete. A minimal standard for morphosyntactic evaluation should include:
- Data provenance documentation: Which languages, genres, and sources were included, and in what proportions?
- Tokenization transparency: Which algorithm was used, and what is the vocabulary size? How are morphological boundaries handled for the target language?
- Paradigm coverage analysis: For the morphosyntactic phenomenon under study, what is the distribution of forms in the training data?
Without these, you can’t separate a model’s inherent capability from the artifacts of its training data. A model that fails on Swahili noun-class agreement may be perfectly capable of learning the system. It just wasn’t given the data to do so.

Practical Takeaways for Model Selection and Analysis
For anyone building or selecting models for morphosyntactically sensitive tasks, a few evidence-backed practices can reduce the risk of data-induced failures:
- Audit the pre-training data. If the data card is incomplete, treat morphosyntactic claims with skepticism. Request language-specific token counts and genre breakdowns.
- Test on typologically diverse probes. A model that performs well on English agreement may fail on ergative alignment or head-final structures. Use diagnostic sets that isolate specific morphosyntactic features.
- Compare models trained on different data mixtures. When possible, evaluate two models with identical architectures but different pre-training corpora. Differences in morphosyntactic accuracy can be attributed to data with higher confidence.
- Watch for tokenization artifacts. If a model consistently fails on inflected forms that share a stem with frequent seen forms, the tokenizer may be splitting the morphemes in a way that obscures the relationship.
FAQ
- Why does pre-training data matter more than model architecture for morphosyntax?
- Architecture sets the upper bound on what can be learned, but pre-training data determines what is actually learned. For most current architectures, the capacity to represent complex morphosyntax exists. The limiting factor is whether the training signal is strong and consistent enough to shape the parameters. Controlled studies that vary data while holding architecture constant consistently show that data composition explains more variance in morphosyntactic accuracy than architectural choices.
- How can I tell if a morphosyntactic error is due to data sparsity or a genuine modeling limitation?
- Check the frequency of the relevant construction in the pre-training data. If the construction appears fewer than a few hundred times, sparsity is the likely culprit. If it appears thousands of times and the model still fails, the issue may be with tokenization, context window, or the model’s inductive biases. Cross-lingual transfer experiments—testing whether the model can handle analogous constructions in a high-resource language—can also help disambiguate.
- Does adding more data always improve morphosyntactic accuracy?
- Not necessarily. Adding data from a typologically different language can introduce interference for rare or parametrically complex features. The composition of the added data matters as much as the quantity. Targeted data augmentation with morphosyntactically diverse examples in the target language is more effective than simply increasing corpus size with unrelated text.
- What is the minimum proportion of a language needed in pre-training data for reliable morphosyntax?
- There is no universal threshold. For high-frequency, regular paradigms like English subject-verb agreement, less than 1% of the total corpus may suffice. For low-frequency or irregular paradigms in morphologically rich languages, 5–10% may be necessary. The key variable is not the language’s overall proportion but the absolute frequency of the specific morphological contrasts the model must learn.
Conclusion: Data-Centric Evaluation as a Path Forward
The morphosyntactic behavior of a model is a fingerprint of its pre-training data. Claims about cross-linguistic capabilities are hollow without a corresponding account of what the model was exposed to. As the field moves toward more rigorous evaluation standards, data transparency must become a baseline requirement, not an afterthought. Future work on this site will examine how data repetition and deduplication affect syntactic generalization, building on the foundation laid here.
This article is part of an ongoing series on empirical evaluation of morphosyntactic generalization. For related discussions, see our analysis of tokenization effects on morphological segmentation and the role of code-switching in pre-training data.