How Models Generate Coherent Narrative Without Event Representation

The Illusion of Narrative Competence

Consider a deceptively simple experiment. You give a large language model a short, carefully constructed narrative: “Elias woke up at dawn. He brewed a pot of coffee and drank a cup while reading the news. At 7:30 AM, he received a worrying email from his boss. Distracted by the email, he lost track of time. Consequently, he missed his 8:15 AM train to the city.” Then you ask the model to summarize the passage. It produces: “Elias overslept and, after drinking his morning coffee, was startled by an email. Consequently, he missed his train.” Grammatically impeccable. A causal connective (“Consequently”) logically links the email to the missed train. The model has identified the protagonist, the inciting incident, and the consequence.

But did the model understand the events? You probe further with a follow-up question that requires temporal reasoning: “Could Elias have drunk his coffee after missing the train?” A human reader, having constructed a mental model of the narrative’s timeline, would immediately recognize this as impossible. The coffee was consumed before the distraction, which occurred before the missed train. The model, however, frequently fails this temporal ordering test. It might answer, “Yes, because drinking coffee and missing a train are independent events,” failing to grasp that the narrative established a specific, unidirectional timeline. The model has successfully summarized the plot without representing the temporal structure of the plot it just summarized.

Push harder. Ask it, “If Elias had not received the email, would he have missed the train?” The model might say, “It is impossible to know,” or it might generate a plausible-sounding but logically unsupported answer. It lacks the causal model to reason about counterfactuals. This failure is not a random error. It is a systematic consequence of the architecture. The model has learned the surface form of narrative without the underlying structure. The coherence in the generated text is a statistical artifact of pretraining data, not evidence of event-level understanding. To see why this happens, we need to examine the difference between generating coherent text and maintaining a propositional model of the world.

What Event Representation Requires

In computational linguistics, event semantics treats an event as a discrete entity with participants, a temporal location, and a causal role. When humans read a narrative, they do not merely process a sequence of words. They construct a mental model of the events—what cognitive psychologists and linguists call a situation model. This model tracks who did what, when it happened relative to other events, where it happened, and how the events are causally related. It allows us to answer questions that require inference, such as whether Elias was holding a coffee cup when he read the email, even if the text never explicitly states it. We infer it from the temporal overlap of the two events in our mental model.

Transformers process text as sequences of tokens. Their objective during pretraining is to predict the next token given the preceding context. This distributional approach captures surface-level correlations beautifully, but it does not naturally build a structured representation of events. The model learns that the phrase “was startled by an email” is frequently followed by words indicating distraction or haste, and that “missed his train” is a plausible consequence of distraction. It learns the narrative grammar of consequence without representing the underlying mechanics of time and causality. The model’s internal state is a high-dimensional embedding of the context, not a symbolic model of events.

The situation model also tracks causal dependencies between events. If Elias missed his train because he was distracted by an email, the human reader understands that removing the email would remove the distraction and potentially prevent the missed train. This counterfactual reasoning is a hallmark of genuine event understanding. The model, lacking an explicit causal model, can only predict that “missed train” is a likely continuation of “distracted by email.” It cannot reason about what would happen if the email had not arrived. This inability to engage in counterfactual reasoning is a direct consequence of the absence of a propositional attitude toward the events. The model does not believe that Elias missed the train. It merely predicts the sequence of words.

For background on how linguistics and philosophy differentiate between surface-level text generation and the underlying propositional structure required for genuine semantic understanding, the foundational concepts of event semantics and discourse coherence provide essential reference points. Event semantics requires binding variables to participants—representing “Elias” as the agent of “drinking” and “missing”—and assigning truth conditions to temporal relations. These operations exceed the pattern-matching capabilities of a purely sequence-to-sequence architecture. The model’s success at narrative summarization is a triumph of syntactic and collocational fluency, not a demonstration of propositional competence. It can generate a summary that sounds like it understands the events, but it cannot manipulate the events themselves.

The Statistical Artifact of Coherence

When a model generates a summary with a coherent plot structure, it is leveraging the immense regularity of human storytelling. Narratives in the training data follow predictable patterns: conflicts escalate, actions have consequences, and temporal sequences are marked by specific linguistic cues. The model learns these cues with high fidelity. It knows that “because” and “therefore” signal causality, and that “then” or “after” signal temporal succession. It learns that a story about a missed train often involves an explanation of why the train was missed.

Using these cues correctly, however, is not the same as understanding the relations they represent. The model’s internal state does not contain a temporal graph where “drink coffee” is a node that precedes the “read email” node. Instead, the model has learned that in texts containing “coffee,” “email,” and “train,” the word “Consequently” is a high-probability transition. This is why models can produce summaries that sound perfectly logical but fail basic consistency probes. The coherence is an artifact of the distributional properties of narrative text, not the product of an internal event model. The model is a master of narrative form. It lacks narrative content.

The distinction matters when evaluating language models for tasks that require genuine reasoning. A model might generate a flawlessly coherent summary of a medical case history, but if it lacks an underlying event representation, it cannot reliably answer questions about whether a symptom appeared before or after a medication was administered. Surface coherence masks a fundamental absence of structured meaning. The model can describe the sequence of events, but it cannot reason about the implications of that sequence. Ask whether the medication could have caused the symptom, and the model will rely on statistical correlations between the medication and the symptom, not on the temporal precedence required for causality.

The Difference Between Coherence and Cohesion

It is worth pausing to distinguish between two concepts often conflated in NLP evaluation: cohesion and coherence. Cohesion refers to the linguistic connections between sentences—pronouns, conjunctions, and lexical chains that tie a text together. Coherence, in the philosophical and linguistic sense, refers to the underlying logical and semantic consistency of the text. A text can be highly cohesive without being coherent, and a text can be coherent without relying heavily on cohesive devices.

Language models are exceptionally good at producing cohesive text. They use pronouns correctly, employ causal connectives, and maintain lexical consistency. But cohesion is a surface phenomenon. Coherence is a deep one. A model can produce a text that is perfectly cohesive but logically incoherent, as in the case of the temporal ordering failure. The model uses “Consequently” to connect the email and the missed train—a cohesive device—but it cannot verify that the consequence is logically consistent with the timeline, which requires coherence.

Most automated metrics—perplexity, n-gram overlap—measure cohesion, not coherence. They reward the model for using the right connectives and maintaining lexical consistency, but they do not check whether the underlying events are logically consistent. To measure coherence, we need tests that probe the model’s understanding of the events, not just its ability to produce cohesive text. The temporal ordering test is a coherence test: it asks whether the model’s internal representation of the events is consistent, not whether it can produce a summary that sounds consistent.

Evaluation Beyond Fluency

The standard evaluation metrics for language generation—BLEU, ROUGE, and even human judgments of fluency—systematically overestimate what models understand because they reward surface-level coherence. If a summary reads well and uses appropriate causal connectives, it is judged as high quality, regardless of whether the model could answer temporal ordering questions about the same narrative. This is a measurement problem. Our evaluation frameworks shape our conclusions about a system’s competence. When we measure narrative quality through fluency and surface coherence alone, we conflate the model’s ability to mimic the form of a story with its ability to represent the content of a story.

Research on how evaluation methodology and framing systematically affect how audiences judge the quality or reliability of information outputs shows that fluency acts as a powerful heuristic for credibility. As noted in methodological research from the Pew Research Center, the way information is presented and evaluated fundamentally shapes perceived competence. In the context of AI, this means a model that sounds coherent is assumed to understand the events it is describing, even when it does not. Human evaluators are particularly susceptible to the illusion of explanatory depth: a fluent, well-structured summary feels like it demonstrates deep understanding, even if the model that generated it cannot answer basic questions about the temporal or causal relations in the narrative.

To accurately assess model understanding, we need evaluation protocols that decouple fluency from propositional accuracy. This means designing tests that require the model to manipulate its internal representation of events, rather than merely generating text about them. Temporal reasoning tests, causal counterfactuals, and coreference probes all serve this function. They force the model to demonstrate whether it has built a situation model or merely learned the statistical regularities of narrative language. A model that can generate a coherent summary but cannot answer the question “Could Elias have drunk his coffee after missing the train?” does not understand the narrative, regardless of how good the summary sounds.

The Architectural Ceiling on Discourse Coherence

True discourse coherence requires more than just connecting sentences with appropriate transitional phrases. It requires tracking speaker commitments, representing causal and temporal relations between events, and maintaining consistent reference across a long context. These capabilities remain architecturally unresolved in transformer-based systems because they require a form of persistent state management that attention mechanisms do not inherently provide.

Attention allows the model to weigh the relevance of different tokens in the context window, but it does not enforce logical consistency. A model can attend to the token “coffee” and the token “train” without maintaining a strict ordering between them. The lack of an explicit symbolic layer for representing time and causality means that the model’s coherence is always probabilistic rather than logical. It can generate a coherent narrative because the distributional probabilities favor coherent sequences, but it cannot guarantee coherence because it has no internal model to check against. This is why long-form generation often degrades into repetition and contradiction: the statistical patterns that maintain coherence over short distances break down over longer contexts where explicit event tracking is required.

This is where the limitations of current architectures become most visible. When human writers or developers use a novel plot generator to map out the causal and temporal dependencies of a narrative, they are engaging in a form of external event representation. The tool acts as a scaffold for human reasoning, helping to externalize the structured relationships that the language model itself lacks internally. The human writer uses the tool to maintain consistency across a long narrative, compensating for the model’s inability to track event relations over extended contexts. The model can suggest what should happen next based on statistical patterns, but the human must verify that the suggestion is consistent with the established event timeline.

The architectural limitation is not a bug that can be patched with more data. It is a fundamental property of systems that process language as sequences of tokens without building explicit models of the situations the language describes. Scaling up the model and the training data improves the statistical approximation of coherence, but it does not bridge the gap between pattern matching and genuine event representation. A larger model will still fail the temporal ordering test, just with more confidence and more fluent rationalizations for its failure.

Practical Takeaways for NLP Practitioners

For practitioners building systems that rely on narrative generation or summarization, the distinction between surface fluency and event representation has immediate practical implications. First, do not assume that a coherent summary implies a coherent understanding. If your application requires temporal reasoning, causal inference, or consistent reference tracking, you must test for these capabilities explicitly, rather than relying on aggregate quality metrics. A model that generates excellent medical summaries may still fail to answer critical questions about the order of events in a patient’s history.

Consider a concrete scenario. You are building a system that generates summaries of legal case histories. The system reads a long narrative of events and produces a summary. If you evaluate the system using BLEU or human judgments of fluency, you might conclude that it is performing well. The summaries will be cohesive and grammatically correct. However, if the system lacks event representation, it may fail to correctly identify the order of events, which is critical in legal reasoning. A plaintiff’s claim may depend on whether they were notified before or after a specific action was taken. If the model cannot represent this temporal relation, it may generate a summary that is cohesive but incoherent—getting the order of events wrong while making it sound plausible.

To avoid this, evaluate the system using a temporal reasoning probe. Provide the system with a case history and ask it to answer questions about the order of events. If the system fails these questions, you know that the summaries it generates are not reliable, even if they score well on fluency metrics. This is the difference between evaluating what the model says and evaluating what the model knows.

Second, consider augmenting neural models with external symbolic representations for events, time, and causality. If the model cannot maintain a situation model internally, you can provide one externally. This is the intuition behind retrieval-augmented generation and neurosymbolic approaches: use the language model for what it does well—generating fluent text—and use a structured database or logic engine for what it does poorly—maintaining consistent event relations. By maintaining an external timeline of events, you can constrain the model’s generation to be consistent with the established narrative.

Third, design evaluation suites that probe for event representation directly. Create minimal pairs of narratives where the surface form is similar but the event structure differs, and test whether the model can distinguish between them. Use temporal ordering probes, causal counterfactuals, and coreference resolution tests as primary metrics of understanding, rather than treating them as secondary benchmarks. The goal is to measure whether the model understands the events, not whether it can describe them fluently. A model that can distinguish between “Elias drank coffee because he missed the train” and “Elias missed the train because he drank coffee” has a deeper understanding than one that can merely generate a summary of either sentence.

Conclusion

The ability to generate a coherent narrative summary is one of the most impressive achievements of modern language models. It is also one of the most misleading. Surface-level fluency masks a fundamental absence of propositional structure, creating the illusion of understanding where there is only statistical pattern matching. As models continue to improve, this gap will become more, not less, visible. The summaries will become more coherent, the narrative flow more smooth, and the failure on temporal reasoning probes more surprising.

Understanding this gap is essential for both AI research and practical deployment. If we measure narrative quality through fluency alone, we systematically overestimate what models know, and we risk deploying systems in high-stakes domains where the difference between describing events and understanding them is critical. The better the model, the clearer the gap between performance and understanding becomes—and the more urgently we need evaluation frameworks that can tell the difference.