Every system that learns from examples ends up absorbing the quiet prejudices of its textbooks. When we watch a model generate text—confidently explaining a legal precedent or spinning a short story—we’re not seeing a flash of original thought. We’re seeing a mirror held up to the vast, messy library it was raised on. The pre-training corpus is that library, and its choice of volumes, the language of its authors, and the years the shelves cover all leave marks that no amount of fine-tuning can fully erase.

The Curriculum Is the Code
Pre-training isn’t a passive scan of words; it’s a deep statistical soak. The model learns to guess the next word in a sentence, and by doing that billions of times, it absorbs the syntax, the common-sense links, and the factual claims scattered through its diet. If the diet is heavy on American English, the model will treat that dialect as the center of the linguistic universe. A phrase like “waiting in line” becomes the default, while “waiting on line” feels like a regional quirk. The behavior isn’t chosen—it’s baked in by the sheer volume of examples.
Think about the temporal slice of the data, too. A corpus cut off in 2019 has no memory of a world shaped by a pandemic. Its understanding of “work” is still anchored to office desks and water-cooler chats. The model isn’t being stubborn when it fails to grasp the nuances of hybrid schedules; it’s simply operating from a fossilized snapshot of human conversation. The data’s timestamp becomes the model’s reality, and any event after that date is science fiction to it.
Linguistic Gravity and Dialect Drift
English exerts a gravitational pull on most large-scale corpora, and that pull warps the model’s behavior in other languages. When a model translates a Japanese novel, it often does so by first mapping the text onto English-shaped concepts, then back out to the target language. The result can be grammatically flawless but culturally hollow—a translation that knows the words but misses the music. Honorifics get flattened, and indirect refusals become blunt “no’s.”
Even within English, the model develops a dialect of its own, a kind of statistical average of everything it’s read. This average skews toward formal, written prose because the web contains more polished articles than casual chats. The model learns to write like a Wikipedia editor, not a friend texting. When we ask it to be conversational, it’s performing a trick it learned from dialogue-heavy novels, not from lived experience. The pre-training mix is a recipe, and the resulting flavor is always a blend of the ingredients.
Social Shadows in the Data
The corpus is a record of human history, and human history is full of uncomfortable patterns. If the training data contains decades of news articles where certain names are linked to certain professions, the model will reproduce those links. It’s not making a judgment; it’s playing the odds. The word “nurse” pulls on “she” because the data shows that pattern millions of times. The word “CEO” pulls on “he” for the same reason. The model is a statistical echo chamber of the past.
This creates a tension between what the model has learned and what we want it to say. We can’t just delete the biased texts; the associations are woven into the fabric of billions of innocent sentences. A medical textbook from the 1970s isn’t overtly sexist, but its pronoun patterns still reinforce old stereotypes. The model’s original tendencies linger like a faint watermark, always ready to reappear if the prompt pushes in the wrong direction.

The Ghosts of Old Books
Many training sets include out-of-copyright books from Project Gutenberg and scanned newspaper archives. These texts are a goldmine for vocabulary and historical language shifts, but they also carry the attitudes of their time. A model raised on 19th-century novels will have deeply encoded the gender roles and colonial assumptions of that era. These aren’t just “facts” the model can list; they’re part of the web of associations that guides its word choices.
Take the word “doctor.” In a modern corpus, it’s gender-neutral in its co-occurrence patterns. In a historical corpus, it’s overwhelmingly masculine. The final model’s behavior is a weighted blend of these patterns, with the balance determined by the ratio of old to new data. The pre-training mix is a dial that controls the model’s social perspective, whether the developers realize they’re turning it or not.
Knowledge Gaps and Confident Fabrication
A model’s expertise in a subject is a direct measure of how much it read about that subject during pre-training. It can write a convincing essay on black holes because it’s seen thousands of physics papers and pop-science articles. It knows the rhythm of the genre, the common analogies, the standard phrases. But ask it about a niche 18th-century poet with only a handful of surviving works, and the model will start to invent. It has no choice; its training objective is to keep generating text, so it pulls on the general patterns of literary biography and creates a plausible-sounding but entirely fictional life story.
This is why a model can summarize a famous play with near-perfect accuracy but completely fabricate the plot of an obscure one. The famous play is dense with quotes, analyses, and forum discussions in the training data. The obscure play might appear only as a title in a list. The model’s behavior shifts from faithful recall to confident confabulation, and the boundary between the two is drawn entirely by the density of the pre-training data.

Code as a Logic Teacher
Mixing code into the pre-training data has a surprising effect: it teaches the model to think in steps. Code is a strange kind of language—unambiguous, rigidly structured, and unforgiving of errors. When a model trains on millions of lines of Python and JavaScript, it learns to follow strict rules and to generate sequences where each token is constrained by a formal grammar. This discipline spills over into natural language tasks, improving the model’s ability to follow instructions, solve math problems, and maintain a logical thread over many paragraphs.
A model trained without code tends to be more fluid and associative, but it struggles with tasks that require holding a chain of reasoning together. The difference isn’t about raw intelligence; it’s about the shape of the training data. Code acts like a mental scaffold, teaching the model to treat a sequence of words as a structured argument rather than a free-association ramble. The percentage of code in the pre-training mix is a direct lever controlling the model’s systematicity.
When the Test Is in the Textbook
A subtle but serious way pre-training data shapes behavior is through contamination. If the training corpus accidentally includes examples from popular evaluation benchmarks, the model’s scores on those benchmarks become meaningless. The model isn’t reasoning through the problem; it’s simply recalling the answer it’s already seen. This creates a false picture of the model’s abilities, because its performance on the benchmark doesn’t predict how it will handle genuinely new problems.
Spotting this contamination is a huge challenge, since it requires checking the entire training corpus against every possible test question. The model’s behavior in these cases is a perfect reflection of its data: it looks like it’s thinking, but it’s really just remembering. The line between reasoning and memorization is one of the deepest questions in the study of these systems, and the answer is hidden in the composition and deduplication of the pre-training data.
FAQ
Why does a model sometimes refuse to answer a question, even if the information is in its training data?
Refusal behavior usually comes from fine-tuning and alignment work done after pre-training, not from the pre-training itself. During pre-training, the model just learns to complete text in a statistically likely way; it has no concept of a harmful or inappropriate answer. The tendency to refuse is taught later, through examples of desired assistant behavior that include politely declining certain queries. Still, the pre-training data plays a role: if the corpus contains very few examples of people politely refusing requests, the model may find it harder to learn this behavior during fine-tuning.
Can you fix a model’s biases by simply removing biased data from the pre-training corpus?
Removing explicitly biased data is a start, but it’s not enough. Bias often hides not in overtly offensive texts but in the statistical co-occurrence patterns across millions of ordinary documents. For example, deleting all texts with sexist language wouldn’t break the link between “nurse” and female pronouns if the remaining corpus still contains a century of medical literature with that statistical pattern. Addressing these deep biases requires more sophisticated techniques, such as data augmentation, counterfactual data generation, or architectural interventions during training, rather than simple filtering.
How does the choice of pre-training data affect a model’s ability to handle multiple languages?
A model’s multilingual ability is almost entirely determined by the language distribution in its pre-training data. It will perform well on languages that are well-represented and poorly on low-resource languages. The type of data matters, too. If the non-English data consists mostly of texts translated from English sources, the model will produce text in that language that carries the stylistic and cultural fingerprints of English. True multilingual fluency requires a pre-training corpus with a rich variety of original, native content in each target language, covering diverse domains and registers.
Why do models sometimes generate text that is factually correct but from a completely wrong time period?
This temporal confusion is a direct result of the model’s training objective. During pre-training, the model isn’t taught to distinguish between a text from 1922 and a text from 2022. It treats all tokens as part of a single, timeless probability distribution. When asked about a current event, it may draw on patterns from historical texts that discuss similar themes, leading to an anachronistic response. The model’s behavior is a blend of all the time periods present in its data, weighted by their frequency, without an intrinsic understanding of chronology unless that understanding is explicitly modeled in the training process.