When a model spits out a sentence with mangled grammar, it’s easy to shrug and call it a glitch. But look closer. The error usually isn’t random—it’s a fossil of the pre-training corpus, a direct reflection of what the model saw, how often it saw it, and in what contexts. This piece digs into that tight coupling: how the composition, curation, and statistical shape of pre-training data constrain morphosyntactic behavior, especially across languages. We’ll walk through data determinism, tokenization artifacts, contamination, and what all of this means for building evaluations that actually tell you something.
Data Determinism and the Generalization Mirage
Call it data determinism: the idea that a model’s grammatical abilities are, to a first approximation, a function of the patterns in its training corpus. When a model nails subject-verb agreement in English but fumbles it in Hindi, the gap isn’t a mystery. It’s a direct consequence of token frequencies, subword splits, and the syntactic contexts that were—or weren’t—present in the data.
This becomes painfully clear with morphologically rich languages. A model might breeze through Spanish gender agreement because the training data was saturated with it. Swap in a language with a smaller digital footprint, like Swahili, and the same model falls apart. It hasn’t learned an abstract rule. It’s learned a surface-level statistical proxy that only holds where the data is dense enough to support it. Cross-linguistic performance isn’t a measure of grammatical competence; it’s a measure of corpus similarity.
Corpus Composition: The Hidden Variable
Too many evaluations treat the model as the only variable that matters, while the pre-training corpus sits in the background like an unexamined assumption. That’s a problem. If two models differ on a benchmark, the difference might have nothing to do with architecture or training objectives. It might just be that one model’s corpus had more relative clauses, or more formal registers, or fewer code-switched sentences.
Take long-distance dependencies. A model trained on a diet of news articles and literary fiction will have seen plenty of nested structures. One trained mostly on conversational text or computer code won’t have. Test both on the same syntactic benchmark, and the first will look smarter. But you’re not measuring grammatical ability—you’re measuring the overlap between the benchmark and the training distribution. Without documenting that overlap, the results are uninterpretable.
Frequency Effects and the Illusion of Learning
Pre-training data doesn’t just determine what a model can do; it determines how robustly it can do it. High-frequency constructions get baked into the model’s weights so thoroughly that performance looks flawless. Low-frequency ones, even if they’re perfectly grammatical, trigger hesitation and errors. The model treats them as edge cases because, statistically, they were.
This creates a sharp performance gradient that mirrors the token frequency distribution of the training corpus. A model might score 95% on a subject-verb agreement task when the subject and verb are adjacent. Insert a prepositional phrase between them—a rarer pattern in the data—and accuracy can crater. The model didn’t learn the rule. It learned a shortcut that breaks when the surface statistics shift.
Tokenization: Where the Trouble Starts
Pre-training data shapes behavior at a level even more fundamental than syntax: the tokenization scheme. Most tokenizers are built for English and a handful of other high-resource languages. They use algorithms like Byte-Pair Encoding (BPE) trained on the same corpus as the model. For morphologically rich languages, this often means chopping a single inflected word into subword pieces that don’t align with actual morphemes.
The fallout is immediate. To produce the correct case form of a Finnish noun, a model might need to generate three or four subword tokens in exactly the right sequence. If the pre-training corpus didn’t contain enough examples of that full paradigm, the model never had a chance to learn the systematic relationship between the pieces. The resulting errors get blamed on the model’s lack of morphological knowledge, but the real culprit is the interaction between tokenization and data sparsity.

Contamination: When Memorization Wears a Mask
Pre-training data can also determine behavior in a more direct and misleading way: by leaking into evaluation benchmarks. When test items or their near-duplicates appear in the training corpus, the model can ace the task through memorization, not generalization. This is a known headache across the field, but its specific impact on morphosyntactic evaluation often gets downplayed.
Morphosyntactic benchmarks frequently rely on templates—fill-in-the-blank with the correct verb form, choose the right case ending. If the pre-training corpus includes similar templated material from language-learning sites or grammar drills, the model may have effectively seen the test items during training. The scores then reflect data overlap, not grammatical ability. Auditing pre-training corpora for benchmark leakage should be standard practice, but in morphosyntax research, it’s still the exception.
Designing Evaluations That Acknowledge the Data
If pre-training data and model behavior are this tightly linked, what does a responsible evaluation look like? Start with transparency. Without knowing the languages, domains, and sources in the training corpus, you can’t interpret morphosyntactic results. A model that shines on a Czech case-marking task might simply have been trained on a mountain of Czech text; its failure on a typologically similar language like Polish may reflect nothing more than data scarcity.
Next, probe the boundaries of data exposure. Instead of asking whether a model can do a task, ask how performance degrades as the task moves away from the training distribution. Vary the lexical items, syntactic structures, or languages in the test set systematically and watch the performance gradient. A model that has genuinely acquired a morphosyntactic generalization should show a gradual decline, not a sudden drop-off.
Controlled Pre-Training and Practical Alternatives
The cleanest way to establish a causal link between data and behavior is through controlled pre-training experiments. Train several models on corpora that differ only in the presence or frequency of a specific construction, then measure the effect on downstream performance. This approach has been used to study filler-gap dependencies and reflexive binding, producing clear evidence that exposure to particular syntactic patterns is necessary for acquisition.
But controlled pre-training is expensive and often impractical. A more accessible alternative is to use probing classifiers and behavioral tests designed to be sensitive to data distribution. A probing classifier trained on model representations can reveal whether a morphosyntactic feature is encoded, but only if the probe’s own training data is carefully controlled to avoid confounding with the pre-training distribution.

Frequently Asked Questions
Why does pre-training data matter more for morphosyntax than for other linguistic phenomena?
Morphosyntax involves systematic, rule-governed patterns that are often sparse in natural corpora. Unlike lexical semantics, where a word’s meaning can be inferred from a few hundred contexts, morphosyntactic rules require exposure to diverse paradigmatic forms. If the pre-training data lacks sufficient examples of a particular inflectional paradigm or syntactic alternation, the model has no basis for acquiring the underlying rule, regardless of its architectural capacity.
Can a model generalize to unseen morphosyntactic constructions?
Generalization to completely unseen constructions is rare and typically limited to cases where the unseen construction is a straightforward analogical extension of a seen one. For example, a model trained on English past-tense forms may correctly produce the past tense of a novel verb if it follows a regular pattern. However, for more complex phenomena like long-distance agreement or case stacking, generalization is minimal unless the pre-training data contains structurally similar examples in some language.
How can I tell if my evaluation results are due to data memorization?
One practical approach is to compare performance on high-frequency versus low-frequency items within the same task. If the model performs well on frequent items but poorly on rare ones, the results are likely driven by memorization. Another method is to test the model on carefully constructed minimal pairs that differ only in the target morphosyntactic feature, ensuring that both members of the pair are equally plausible in terms of surface statistics. A model that relies on shallow heuristics will struggle to distinguish them.
What steps can researchers take to mitigate data determinism in evaluations?
Researchers should document the pre-training data composition as thoroughly as possible and report performance broken down by data frequency bins. When using public benchmarks, they should check for benchmark contamination and consider constructing adversarial test sets that specifically probe the limits of the training distribution. Finally, they should complement behavioral testing with causal analyses, such as ablation studies or controlled pre-training, to establish whether a model’s success on a task is truly due to grammatical competence.

Where We Go From Here
The relationship between pre-training data and model behavior isn’t a nuisance variable to be controlled away. It’s the central object of study for anyone serious about the empirical foundations of morphosyntactic competence. Treating data composition as a first-class variable lets us move past binary pass/fail benchmarks and toward a clearer picture of how statistical learning interacts with linguistic structure.
This article is part of a wider look at evaluation methodology. Upcoming pieces will tackle the role of tokenization in cross-linguistic transfer, the design of minimal-pair tests for syntactic phenomena, and the challenges of building balanced multilingual benchmarks. Each of these threads circles back to the same core principle: the data makes the model, and understanding the model means understanding the data.