Pre-training data is the initial, large-scale text corpus used to shape a model’s statistical expectations before any task-specific adjustment. In morphosyntax research, this corpus is not a neutral substrate. It encodes frequency distributions, register biases, orthographic conventions, and cross-linguistic asymmetries that later surface as model behavior. Adjacent concepts include data mixture ratios, tokenization artifacts, corpus provenance, and distributional priors. For researchers who evaluate language model failures in morphosyntax and cross-linguistic generalization, understanding pre-training data is a prerequisite for interpreting any downstream result. A model’s grammatical judgments, agreement errors, and case-marking choices are often less about architecture and more about what the model was exposed to during pre-training.
This article examines the mechanisms through which pre-training data determines model behavior. It focuses on measurable effects: frequency effects in morphology, cross-linguistic coverage gaps, register and domain skew, and the consequences of data filtering. The goal is to provide a framework for reading model outputs as artifacts of corpus design rather than as evidence of abstract linguistic competence.
What Pre-Training Data Actually Controls
Pre-training data controls the distributional evidence a model uses to learn form-function mappings. In morphosyntax, this includes agreement patterns, word order regularities, case assignment, and the interaction between morphology and discourse structure. When a model produces a nonstandard agreement form, the first question should be whether the form was underrepresented, absent, or systematically filtered from the training corpus.
Three properties of pre-training data have the strongest influence on model behavior:
- Token frequency: High-frequency morphological forms are learned earlier and reproduced more reliably. Low-frequency forms, even when grammatical, are more likely to be replaced by a dominant alternative.
- Context diversity: A form that appears in many syntactic contexts generalizes better than a form that appears frequently but only in one construction.
- Cross-linguistic balance: Languages with limited representation in the corpus produce weaker and less stable morphosyntactic behavior, independent of the model’s total parameter count.
These properties are not independent. A language can have high token counts but low context diversity, or broad coverage but shallow morphological annotation. Researchers who evaluate model behavior without inspecting these corpus properties risk attributing corpus artifacts to model capacity.
Frequency Effects in Morphosyntactic Production
Frequency effects are among the most reproducible findings in language model evaluation. Models trained on skewed corpora tend to overproduce frequent forms and underproduce rare but grammatical alternatives. This is especially visible in inflectional morphology.
Consider subject-verb agreement in English. The third-person singular present tense marker -s is frequent in written corpora, but its distribution is uneven across registers. In corpora dominated by news text, the form appears regularly. In corpora dominated by informal dialogue, it competes with unmarked forms. A model pre-trained on a news-heavy corpus may overapply the -s marker in contexts where a human speaker would not. The model is not making a grammatical error in the traditional sense; it is reproducing the statistical center of its training distribution.
Similar effects appear in languages with richer inflectional systems. In Russian, case endings vary by declension class, gender, and animacy. A model trained on a corpus with disproportionate representation of one declension class will favor that class’s endings even when the context requires another. The error is systematic, not random, and it traces directly to corpus composition.
Measuring Frequency Effects
Researchers can measure frequency effects by comparing model outputs against corpus frequency tables. A simple approach is to extract the top morphological variants for a given lemma from the pre-training corpus and compare them with the model’s preferred variants in controlled prompts. When the model’s preferences align with corpus frequencies, the behavior is likely distributional rather than rule-based.
This method has limits. Corpus frequency tables are not always available for proprietary pre-training data. In those cases, researchers can use proxy corpora that approximate the likely training distribution. The results are less precise but still informative for identifying broad frequency biases.
Cross-Linguistic Coverage and Generalization
Cross-linguistic generalization is a central concern in morphosyntax evaluation. A model that performs well on English agreement may fail on Swahili noun class agreement or Basque ergative marking. The failure is often not a sign of architectural limitation but of corpus imbalance.
Pre-training corpora are heavily skewed toward a small number of high-resource languages. English, Chinese, Spanish, and a few others dominate the token distribution. Many languages appear only in small quantities, often from limited domains such as religious texts, legal documents, or web crawls with inconsistent quality. This skew has direct consequences for morphosyntactic behavior.
For example, a model trained on a corpus where a language appears mostly in translated text may inherit translationese patterns. Word order may be influenced by the source language, and morphological marking may be simplified or regularized. The model then reproduces these patterns in evaluation tasks, creating the appearance of grammatical failure when the underlying issue is corpus provenance.
Data Mixture Ratios
Data mixture ratios determine how much weight each language receives during pre-training. These ratios are often set by heuristics or availability rather than linguistic criteria. A language with 0.1% of the training tokens cannot be expected to exhibit the same morphosyntactic stability as a language with 10%.
Researchers should treat data mixture ratios as an experimental variable. When comparing models, differences in language performance may reflect differences in mixture ratios rather than differences in model quality. Reporting these ratios, when known, is essential for reproducible evaluation.
Register and Domain Skew
Pre-training corpora are not uniform in register. They mix news articles, web forums, academic papers, fiction, and other text types in proportions that rarely match natural language use. This skew affects morphosyntactic behavior in predictable ways.
News text tends to favor complete sentences, explicit subjects, and standard punctuation. Web forum text favors fragments, ellipsis, and nonstandard orthography. Academic text favors nominalizations and complex embedding. A model pre-trained on a corpus dominated by one register will reproduce that register’s morphosyntactic patterns even when the prompt suggests another.
This is especially relevant for evaluation tasks that use controlled prompts. If the prompt is written in a formal register but the model’s training data is dominated by informal text, the model may produce informal morphosyntax. The mismatch is not a failure of grammatical knowledge but a register prior imposed by the corpus.
Domain-Specific Morphosyntax
Domain skew also affects domain-specific morphosyntax. Legal text, for example, uses distinctive modal constructions and passive forms. Medical text uses specialized nominal compounds. A model trained on general web text may not have enough exposure to these patterns to reproduce them reliably.
For morphosyntax researchers, this means that evaluation tasks should specify the register and domain of the expected output. A model’s performance on a legal-domain agreement task cannot be interpreted without knowing whether legal text was represented in the pre-training corpus.
Data Filtering and Its Consequences
Pre-training corpora are rarely used in raw form. They are filtered to remove duplicates, low-quality text, and content deemed inappropriate. These filtering steps can have unintended effects on morphosyntactic behavior.
Duplicate removal, for example, reduces the frequency of repeated phrases and formulaic expressions. This can lower the model’s exposure to fixed morphosyntactic patterns, making it less likely to reproduce them. Quality filtering, which often targets short or ungrammatical text, can remove exactly the kind of nonstandard morphosyntax that researchers may want to study.
Filtering can also introduce biases. If a filter removes text with nonstandard orthography, it may disproportionately remove text from certain dialects or languages. The resulting corpus is cleaner but less representative. Models trained on filtered corpora may then underperform on dialectal morphosyntax, not because the architecture cannot handle it, but because the training data was systematically reduced.
Tokenization as a Pre-Training Decision
Tokenization is often treated as a technical detail, but it is a pre-training decision with direct morphosyntactic consequences. The choice of tokenizer determines how words are split into subword units, which affects how morphological information is represented.
For morphologically rich languages, subword tokenization can split inflectional affixes from stems. This can help the model generalize across inflected forms, but it can also obscure the relationship between a stem and its affixes if the split is inconsistent. A tokenizer that treats a case ending as part of the stem in one context and as a separate token in another creates inconsistent evidence for the model.
Tokenization also interacts with corpus frequency. Rare morphological forms may be split into many subword tokens, while frequent forms are kept whole. This creates an asymmetry in how the model learns different forms, independent of their grammatical status.
Evaluating Model Behavior Through Corpus Lenses
A practical approach to evaluating model behavior is to treat every output as a corpus artifact until proven otherwise. This means asking a series of questions before attributing a behavior to linguistic competence:
- Was the relevant form present in the pre-training corpus, and at what frequency?
- Was the form distributed across multiple contexts or concentrated in one domain?
- Was the language represented in sufficient quantity to support stable morphosyntactic learning?
- Did filtering or tokenization alter the distribution of the form?
These questions are not always answerable, especially for proprietary models. But they provide a disciplined framework for interpreting results. When the answers are unknown, researchers should state that uncertainty explicitly rather than assuming the model’s behavior reflects abstract grammatical knowledge.
Case Study: Agreement Errors in Low-Resource Languages
Consider a model that produces incorrect subject-verb agreement in a low-resource language. The error could be caused by several pre-training factors: insufficient token counts, limited context diversity, register skew, or filtering artifacts. Without corpus inspection, the error is uninterpretable.
A more informative approach is to compare the model’s errors with the most frequent forms in the available corpus. If the model consistently produces the most frequent form regardless of context, the behavior is likely a frequency effect. If the model produces forms that do not appear in the corpus at all, the behavior may reflect a different mechanism, such as cross-linguistic transfer from a related language.
This kind of analysis requires access to corpus statistics, which are not always published. Researchers can build their own frequency tables from publicly available corpora that approximate the training distribution. The results will be approximate, but they can still distinguish between frequency-driven and transfer-driven errors.
Practical Takeaways for Morphosyntax Evaluation
For researchers who evaluate language model failures in morphosyntax, the following practices can improve the interpretability of results:
- Report corpus context: When possible, document the pre-training corpus composition, including language distribution, register mix, and filtering steps.
- Use frequency-matched controls: Compare model outputs against corpus frequency tables to identify distributional effects.
- Test across registers: Run the same morphosyntactic task in multiple registers to separate register priors from grammatical knowledge.
- Include low-resource languages: Evaluate cross-linguistic generalization with languages that vary in corpus representation, not just high-resource languages.
- Treat tokenization as a variable: When comparing models, note differences in tokenization that may affect morphological representation.
These practices do not require access to proprietary training data. They require a shift in interpretive stance: from asking whether a model is grammatical to asking what distributional evidence shaped its behavior.
Limitations and Open Questions
The relationship between pre-training data and model behavior is not fully deterministic. Models can generalize beyond their training distribution in some cases, and architecture can modulate the effects of corpus skew. The same corpus can produce different behaviors in different models, depending on optimization, capacity, and training duration.
There are also open questions about the role of data order. Pre-training data is presented in a specific sequence, and the order can affect what the model learns and when. This is less studied than corpus composition, but it may explain some inconsistencies in model behavior.
Finally, the field lacks standardized methods for reporting pre-training data properties. Without such standards, comparisons across models remain difficult. A shared vocabulary for describing corpus composition, filtering, and tokenization would improve the reproducibility of morphosyntax evaluation.
FAQ
Why do language models make different morphosyntactic errors in different languages?
The errors often reflect differences in pre-training data coverage. Languages with limited representation in the training corpus provide less distributional evidence for morphological patterns, leading to less stable behavior. Register skew and filtering can further reduce the quality of the evidence for specific languages.
Can a model learn morphosyntax without explicit grammatical rules?
Yes. Models learn morphosyntactic patterns from distributional regularities in the pre-training corpus. They do not need explicit rules to reproduce agreement, case marking, or word order. However, the patterns they learn are shaped by the corpus, including its biases and gaps.
How can researchers evaluate model behavior without access to proprietary pre-training data?
Researchers can use publicly available corpora that approximate the likely training distribution. They can also compare model outputs across registers and languages to identify distributional effects. While these methods are less precise than direct corpus inspection, they can still reveal frequency-driven and transfer-driven behaviors.
Does more pre-training data always improve morphosyntactic performance?
Not necessarily. More data can improve performance if it increases context diversity and cross-linguistic coverage. But if the additional data is skewed toward a single register or language, it can reinforce existing biases. The composition of the data matters as much as the quantity.
Next Steps for This Publication
This article establishes a framework for reading model behavior as a product of pre-training data. A natural follow-up is a detailed examination of tokenization effects on morphological generalization, with examples from agglutinative and fusional languages. Another direction is a comparative study of register skew in publicly available corpora and its consequences for agreement evaluation. Both topics extend the current analysis and deepen the site’s coverage of evaluation methodology.
Readers with specific corpus examples or evaluation results are invited to share them. Concrete cases of frequency-driven errors or cross-linguistic transfer help refine the framework and make it more useful for the research community.


