When a language system generates text, it isn’t reaching for some pristine, universal grammar. It’s echoing the statistical grooves worn into its training data. The makeup of that initial corpus—what’s kept, what’s discarded, and how heavily each piece is weighted—shapes everything: the vocabulary it reaches for, the facts it can retrieve, the styles it can …
Continue reading What a Model Learns Before It Learns Anything Else: The Hidden Influence of Training Data
How Models Generate Coherent Narrative Without Event Representation
The Illusion of Narrative Competence Consider a deceptively simple experiment. You give a large language model a short, carefully constructed narrative: “Elias woke up at dawn. He brewed a pot of coffee and drank a cup while reading the news. At 7:30 AM, he received a worrying email from his boss. Distracted by the email, …
Continue reading How Models Generate Coherent Narrative Without Event Representation
Why Screenplay Generation Exposes the Absence of Propositional Commitment in Language Models
Picture a scene. A character named Vera mentions a locked door in the basement. She tells her companion not to go near it. Cut to a diner. Cut back to the house — but the narrative has moved on to a phone call. By scene twelve, someone opens the door. In a competently written screenplay, …
Continue reading Why Screenplay Generation Exposes the Absence of Propositional Commitment in Language Models
The Ghost in the Dataset: How Pre-Training Data Shapes Model Behavior
When you watch a model generate text, you’re not seeing pure reasoning. You’re catching a reflection—often a distorted one—of the vast textual world it absorbed during its earliest training. The pre-training corpus is the model’s foundational diet, and its composition (the biases, the blind spots, the stylistic tics) directly sculpts how the system behaves. A …
Continue reading The Ghost in the Dataset: How Pre-Training Data Shapes Model Behavior
The Hidden Curriculum: How Training Data Shapes a Model’s Mind
Every system that learns from examples ends up absorbing the quiet prejudices of its textbooks. When we watch a model generate text—confidently explaining a legal precedent or spinning a short story—we’re not seeing a flash of original thought. We’re seeing a mirror held up to the vast, messy library it was raised on. The pre-training …
Continue reading The Hidden Curriculum: How Training Data Shapes a Model’s Mind