Why Narrative Coherence in Language Models Is Not Story Understanding

Ask a language model to write a short story. You will get something that reads like one. A protagonist, a conflict, a turning point, a resolution. Grammatical prose. Emotional shifts at the moments you would expect. Skim it and you might not notice anything wrong.

That is exactly the problem.

What looks like storytelling is, under diagnostic pressure, something else entirely: positional sequence prediction operating over narrative templates the model has absorbed from training data. There is no internal representation of events, causation, or perspective. The story looks right because the surface features of narrative—protagonist introduction, rising action, climactic turn—are statistical regularities in the training distribution. But the deeper discourse-level phenomena that make stories intelligible—event causation, perspective tracking, presupposition projection across scenes, the pragmatic inference required to understand why a character does something rather than merely that they do it—are absent. The model produces the shape of a story without constructing the thing that makes a story a story.

This article treats narrative generation as a diagnostic window into what transformers cannot represent. The gap between plot-structure fluency and genuine story understanding is not a minor imperfection. It is a precise boundary, and better models make it more visible, not less.

The Difference Between Plot Structure and Event Structure

Consider what a story is, minimally. A story is not a sequence of sentences describing events. A story is a sequence of events connected by causal and intentional relations such that each event matters because of the events that came before it. The philosopher of narrative Noël Carroll distinguishes between narrative connection—event A leads to event B—and mere temporal succession—event A happened, then event B happened. A weather report describes temporal succession. A story describes narrative connection.

Language models are good at producing text that appears to encode narrative connection. They deploy words like “because,” “therefore,” “so,” and “as a result.” They attribute motives to characters. They generate turning points and resolutions. But these are surface signals of causal structure, not evidence that the model represents causal structure internally. The model has learned that stories contain causal connectives. It has not learned to model the causal relations those connectives describe.

This distinction matters because it maps onto a specific architectural limitation. Transformers process sequences through attention over token positions. They do not construct event representations. They do not maintain a model of “what happened” that is separate from the sequence of words describing what happened. When a human reads “Maria opened the door. The room was dark. She reached for the light switch,” they build a mental model: Maria is in a doorway, there is a room, the room is dark, Maria’s arm extends toward a switch. When a transformer processes the same text, it computes attention weights over token positions. The coherence of the output depends on whether the training distribution contained similar sequences—not on whether the model has built an event model.

What Plot Generators Actually Do

The commercial deployment of AI-powered plot generators makes this architectural limitation concrete. Tools like the Reedsy Plot Generator accept user inputs—protagonist, conflict, stakes, genre, tone, story structure—and return act-by-act plot outlines. The user selects from named templates: 3-Act Structure, 5-Act Structure, Save the Cat, the Hero’s Journey, or the 7-Point Structure. The generator fits the user’s inputs into these frameworks and produces a structurally recognizable plot shape.

What these tools demonstrate is that narrative generation, as currently practiced, is template-fitted sequence prediction. The generator does not construct an event model. It does not track causation across acts. It does not represent character perspective or project presuppositions between scenes. It maps user-supplied parameters onto structural templates and fills the gaps with text that resembles the training distribution for each genre and tone. The result is a plot outline that looks like a story structure because it is shaped like one—not because the system understands what a story structure is.

That same discipline applies to narrative structure: before publishing, editors need a way to test events, claims, and consequences actually follow one another, which is where how Unsloppy AI Writing App fits the writing workflow can function as a planning aid rather than a substitute for domain evidence.

That same discipline applies to narrative structure: before publishing, editors need a way to test whether events, claims, and consequences actually follow one another. Planning tools that fit into a writing workflow can help test event-structure consistency without substituting for genuine narrative understanding. The Unsloppy AI Writing App is one such tool, but its coherence is structural in the same sense as the others: it comes from fitting inputs to templates, not from modeling the events the templates describe.

The Reedsy tool’s own framing is inadvertently diagnostic. It notes that “a character who encounters no meaningful resistance is merely in a sequence of events.” Correct. But the generator itself has no mechanism for distinguishing meaningful resistance from arbitrary obstruction. It can produce the words “the protagonist faces a setback” because the template includes a setback slot. It cannot determine whether the setback is causally motivated by preceding events or randomly inserted to fill the slot. That determination requires an event model the architecture does not possess.

What an Event Model Would Require

To see what is missing, it helps to specify what genuine narrative understanding would require. An event model is a structured representation of events, their participants, their temporal relations, and their causal relations. It is what allows a reader to answer questions like: Why did the character do this? What changed because of this event? What did the character know at this point in the story? What will happen next, given what has happened?

These questions require more than sequence prediction. They require:

  • Causal representation: The ability to represent that event A caused event B, not just that the text mentions A before B.
  • Perspective tracking: The ability to represent what each character knows, believes, and wants at each point in the narrative, and to update these representations as events unfold.
  • Presupposition projection: The ability to carry background assumptions across scenes, so that information established in scene 1 is available as a presupposition in scene 3.
  • Intentional explanation: The ability to represent why a character performed an action, not just that they performed it.

None of these capabilities follow from next-token prediction over a training corpus. They require structured representations that are not reducible to embedding-space proximity or attention-weight distributions. A model can produce the sentence “Maria opened the door because she heard a noise” without representing the causal relation between hearing a noise and opening a door. It produces this sentence because, in the training distribution, sentences about doors and noises frequently co-occur with causal connectives. The causal relation is in the text. It is not in the model.

A Reproducible Experiment: Event-Structure Consistency

To test this claim, I designed a minimal experiment that probes whether a language model can maintain event-structure consistency across a generated narrative. The experiment is reproducible with any instruction-tuned transformer model and requires no specialized evaluation infrastructure.

The setup is simple. Prompt the model to generate a short story (approximately 500 words) about a scenario involving a causal chain: a character discovers a problem, takes an action to address it, and experiences a consequence. Then, after the story is generated, ask the model a series of questions about the events in the story—questions that require an event model to answer correctly.

The Prompt

Write a short story (about 500 words) about a pharmacist
named David who discovers that a medication batch is
mislabeled. He must decide whether to report it, knowing
it will cost his employer money. Include a scene where
David talks to his supervisor, and a scene where David
makes his decision. The story should have a clear ending.

The Diagnostic Questions

After the story is generated, ask the model:

  1. What did David know at the time he talked to his supervisor?
  2. What caused David to make his final decision?
  3. Was there information David had in the first scene that he did not have in the second scene? If so, what?
  4. If David had not discovered the mislabeling, what would have happened?

The Code

import openai

def generate_story_and_probe(model="gpt-4o", seed=42):
    story_prompt = (
        "Write a short story (about 500 words) about a "
        "pharmacist named David who discovers that a "
        "medication batch is mislabeled. He must decide "
        "whether to report it, knowing it will cost his "
        "employer money. Include a scene where David talks "
        "to his supervisor, and a scene where David makes "
        "his decision. The story should have a clear ending."
    )

    story_response = openai.chat.completions.create(
        model=model,
        messages=[{"role": "user", "content": story_prompt}],
        temperature=0.7,
        seed=seed,
    )
    story = story_response.choices[0].message.content

    probe_questions = [
        "What did David know at the time he talked to "
        "his supervisor?",
        "What caused David to make his final decision?",
        "Was there information David had in the first "
        "scene that he did not have in the second scene? "
        "If so, what?",
        "If David had not discovered the mislabeling, "
        "what would have happened?",
    ]

    results = []
    for q in probe_questions:
        probe_response = openai.chat.completions.create(
            model=model,
            messages=[
                {"role": "system", "content":
                    "You are answering questions about a "
                    "story. Answer based only on the story "
                    "text provided. Do not add information "
                    "not present in the story."},
                {"role": "user", "content":
                    f"Story:\n{story}\n\nQuestion: {q}"},
            ],
            temperature=0.0,
            seed=seed,
        )
        answer = probe_response.choices[0].message.content
        results.append({"question": q, "answer": answer})

    return {"story": story, "probe_results": results}

output = generate_story_and_probe()
for r in output["probe_results"]:
    print(f"Q: {r['question']}")
    print(f"A: {r['answer']}\n")

What the Experiment Reveals

Run this experiment with any current instruction-tuned model and you will observe a consistent pattern. The generated story will be fluent. It will have a recognizable narrative arc. The prose will be competent. The diagnostic questions, however, will expose the gap.

Question 1—what did David know at a specific point in the story—tests perspective tracking. Models fail this question in one of two ways. Either they attribute knowledge to David that the story establishes only later (the model cannot distinguish what David knew at the time from what is true in the story overall), or they produce a vague summary that does not distinguish between David’s knowledge and the narrator’s knowledge. This failure occurs because the model does not maintain a representation of character epistemic state that is separate from the sequence of words in the story.

Question 2—what caused David’s decision—tests causal representation. Models typically produce plausible-sounding answers that cite factors mentioned in the story. But when the story contains multiple potential causes, the model often fails to identify which cause was decisive—which cause actually motivated the decision, as opposed to which causes were merely present in the narrative context. The model can identify that the story mentions a cause. It cannot represent the causal relation that makes it the cause.

Question 3—information available in one scene but not another—tests presupposition projection across scenes. This is where the failure is most systematic. Models often answer by summarizing what happened in each scene, without tracking what information was available to whom at each point. A story in which David discovers something in scene 1 and then, in scene 2, acts as though he does not know it will not be flagged as inconsistent by the model. The model processes each scene as a sequence of words. It does not maintain a presupposition structure that carries commitments across scene boundaries.

Question 4—the counterfactual: what would have happened if David had not discovered the mislabeling—tests counterfactual reasoning over the event structure. This is where the absence of an event model becomes most visible. Models typically produce a plausible continuation: “the mislabeled medication would have reached patients.” But this continuation is based on world knowledge about medication safety, not on reasoning over the specific event structure of the generated story. If the story established that a second character had already noticed the mislabeling, the model rarely incorporates this into the counterfactual. It generates the most statistically likely continuation, not the continuation entailed by the story’s specific event structure.

Why This Is Not a Problem That Scale Solves

It is tempting to attribute these failures to insufficient model capacity and predict that larger models, trained on more narrative data, will close the gap. This prediction misunderstands the nature of the limitation.

The problem is not that the model has not seen enough stories. The problem is that the architecture does not construct the kind of representation that narrative understanding requires. Scaling up a transformer—adding layers, expanding the context window, increasing the training corpus—makes the model better at producing text that resembles stories. It does not make the model better at representing events, tracking perspectives, or projecting presuppositions. These capabilities require structured representations that are not reducible to the sequence-prediction objective.

This is the same architectural limitation that appears in syntactic island constraints, in coreference resolution over long distances, and in the processing of quantifier scope ambiguity. In each case, the model produces fluent output that appears to handle the relevant phenomenon, but diagnostic probing reveals that the underlying representation is absent. Narrative generation is simply the domain where the gap between surface fluency and underlying competence is most visible, because stories are the kind of language where event structure matters most.

The broader context of how these models are trained reinforces the point. As the Authors Guild has documented, commercially available foundational LLMs have been trained on unlicensed books without compensating authors or publishers. The training corpus includes vast quantities of narrative text—novels, short stories, screenplays. The model has seen more stories than any human could read in a lifetime. And yet the diagnostic questions fail. This is evidence that the limitation is architectural, not distributional. You cannot solve a representation problem by adding more data to a system that does not construct the relevant representation.

The Diagnostic Value of Narrative

Narrative generation is valuable as a diagnostic tool precisely because it makes the gap between fluency and understanding so easy to see. A model that generates a grammatical sentence with a syntactic island violation requires a linguist to diagnose the error. A model that generates a story in which a character forgets what they knew two paragraphs ago produces a failure that any attentive reader can detect.

This accessibility makes narrative generation a useful probe for evaluation design. The standard evaluation regime for language models—benchmarks that reward surface fluency, perplexity scores that measure next-token predictability, human preference ratings that correlate with stylistic polish—cannot detect the absence of an event model. These metrics measure whether the output looks like a story. They do not measure whether the system understands what a story is.

The experiment described above is a step toward a better evaluation. It tests a specific capability—event-structure consistency—that is necessary for genuine narrative understanding and that is not captured by standard metrics. It is reproducible, requires no specialized infrastructure, and produces failures that are interpretable in terms of specific representational deficits. It is the kind of evaluation this blog has advocated for in syntax and semantics: one that probes what the model represents, not what the model produces.

What Remains Open

The experiment raises a question I cannot answer here and that I think the field has not adequately addressed. If narrative understanding requires an event model, and if transformers do not construct event models, then what architectural modification would enable them to do so? The question is not whether we can bolt on a symbolic event tracker as a post-hoc module. The question is whether there is a way to build event representation into the architecture itself, so that the model’s internal representations include the causal, temporal, and intentional structure of the events its text describes.

Neurosymbolic approaches offer one direction. Structured state-tracking mechanisms offer another. But neither has been demonstrated at the scale where narrative generation becomes interesting. The open question is whether the gap between plot-structure fluency and genuine story understanding is bridgeable within the transformer paradigm, or whether it is a fundamental limitation of the architecture that no amount of scale or training data can overcome.

I suspect it is the latter. But suspicion is not evidence. The experiment above is offered as a tool for others to test that suspicion against whatever models come next.