Introduction: The Data Diet of a Language Model When a model handles English syntax with ease but stumbles on a simple German relative clause, the problem usually isn’t the architecture. It’s the training data. The mix of languages, the types of texts, and the presence or absence of specific grammatical patterns in the pre-training corpus …
Continue reading How Pre-Training Data Shapes Model Behavior: A Cross-Linguistic View
Author:Terri Lane
How Pre-Training Data Shapes Language Model Behavior
How Pre-Training Data Shapes Language Model Behavior When a language model generates text, it draws on patterns absorbed during pre-training—the massive, initial phase where the model ingests a curated collection of documents. The composition of that collection, from the sources chosen to the proportions of different text types and the decisions about what to exclude, …
Continue reading How Pre-Training Data Shapes Language Model Behavior
Why Discourse Coherence Remains the Hardest Problem in NLP — and What External Scaffolding Can and Cannot Do About It
Why Discourse Coherence Remains the Hardest Problem in NLP — and What External Scaffolding Can and Cannot Do About It Consider two paragraphs: Maria set the ancient vase on the wobbly table. It had been in her family for six generations. The wood groaned under the weight. The detective examined the scene. Maria had told …
Continue reading Why Discourse Coherence Remains the Hardest Problem in NLP — and What External Scaffolding Can and Cannot Do About It
How Pre-Training Data Shapes Language Understanding: A Computational Linguist’s View
Every language system starts with a corpus—a pile of texts that is, for all practical purposes, its only window into human communication. What goes into that pile matters enormously. Parliamentary transcripts, recipe blogs, old novels, forum rants: the mix determines which syntactic patterns stick, which semantic associations form, and which pragmatic conventions the system ends …
Continue reading How Pre-Training Data Shapes Language Understanding: A Computational Linguist’s View
How Pre-Training Data Shapes Model Behavior: A Measured Look at Corpus Effects
Every language model starts not with a clever bit of code, but with a library. The collection of texts—the pre-training corpus—is the model’s only window into human language. Before any fine-tuning, before any reinforcement learning from human feedback, the model is a statistical mirror of that corpus. It learns which words tend to follow others, …
Continue reading How Pre-Training Data Shapes Model Behavior: A Measured Look at Corpus Effects