Why Screenplay Generation Exposes the Absence of Propositional Commitment in Language Models

Picture a scene. A character named Vera mentions a locked door in the basement. She tells her companion not to go near it. Cut to a diner. Cut back to the house — but the narrative has moved on to a phone call. By scene twelve, someone opens the door. In a competently written screenplay, that door was never forgotten. It was a promise, and the script either keeps it or makes its violation meaningful. In a model-generated screenplay, the door is frequently forgotten, or worse, remembered in a way that contradicts the original framing. The wrong character opens it. It migrates to a different floor. Someone references it as though a different person had mentioned it. This is not a coherence problem in the usual sense. The sentences are fluent. The scene headings are correctly formatted. The problem is referential: the model has no mechanism for maintaining a commitment state across turns, because it has no representation of speaker commitment at all.

Screenwriting turns out to be an unusually precise diagnostic for this limitation. A screenplay is not just a long text. It is a chain of scene-level promises, each establishing spatial, temporal, and causal commitments that subsequent scenes must honor or deliberately subvert. The format itself encodes these commitments. Scene headings pin down location and time. Character introductions create entities that persist across pages. Action lines establish physical states that later scenes can alter or reference. As StudioBinder’s guide to how to write a movie script like professional screenwriters makes clear, the structural conventions of screenplay format — from scene headings that anchor spatial-temporal geography to the one-page-equals-one-minute production ratio — are not cosmetic preferences. They are referential scaffolding. Each scene heading is a commitment the rest of the script answers to. When a model generates a screenplay, it must maintain this accountability across thousands of tokens, and that is precisely where the architecture begins to show its limits.

The failure is subtle because it rarely manifests as obvious incoherence. The model produces well-formed sentences. It uses correct slug lines. The dialogue sounds like dialogue. What breaks is the chain of obligations a screenplay weaves: a gun placed in a desk drawer in scene four should be available, absent, or meaningfully unmentioned by scene fifteen. Models frequently lose track of these obligations, not because they cannot attend to distant context, but because they have no representation of what it means to be obligated by a previous statement. The desk drawer mention is, for the model, a token sequence that contributed to a probability distribution at one point in generation. It is not a commitment the model carries forward as a constraint on future output.

Locutionary Production Without Illocutionary Force

Speech act theory, as developed by Austin and Searle, distinguishes between the locutionary act — producing a grammatical utterance with a certain sense and reference — and the illocutionary act, which is what the speaker does in producing that utterance: asserting, promising, warning, requesting, committing. When Vera says ‘don’t go near that door,’ the locutionary content is a prohibition. The illocutionary force is a warning backed by an implied commitment: Vera is committed to the proposition that the door is dangerous, or important, or relevant in a way the companion should respect. A human writer maintaining this scene carries both the content and the force forward. The door is not just a phrase. It is a committed proposition that subsequent scenes must engage with.

Transformers produce locutionary output. They generate text that has the surface characteristics of illocutionary acts — warnings that look like warnings, promises that look like promises — but the architecture contains no representation of the force behind these acts. A sentence in a generated screenplay is a continuation of a token sequence, not a committed proposition. This is why a model can write a scene where a character solemnly vows to protect a secret, and then produce a later scene where the same character casually mentions the secret to a stranger, with no awareness that the earlier vow created an obligation. The model has no commitment state to violate. It never promised anything. It produced tokens that, in training data, tended to follow certain patterns.

This distinction matters because most evaluation frameworks do not test for it. BLEU, ROUGE, and their descendants measure surface overlap. Embedding-based coherence metrics measure semantic similarity between passages. Even human evaluation protocols often focus on local fluency and plausibility rather than tracking whether specific propositional commitments established early in a document are honored late in it. A screenplay evaluation that asks ‘does scene twelve feel consistent with scene three?’ will miss the specific failure: scene twelve contradicts a promise made in scene three, but in a way that feels locally plausible because the model has generated text that resembles the kind of scene where such a contradiction would not occur.

The Diagnostic Value of Long-Form Referential Dependencies

Screenplays are useful precisely because they make referential dependencies explicit and consequential. In casual conversation, a model can drift away from a topic without the user noticing. In a short story, inconsistencies can be absorbed into the general ambiguity of literary prose. A screenplay, however, is a structured document where every element — a character’s name, a prop, a location, a stated intention — creates a specific referential anchor that the production team will physically instantiate. If scene three establishes that the door is in the basement and scene twelve places it on the second floor, that is not an artistic choice. It is a continuity error that makes the script unusable.

This is why an AI script writing app, regardless of its engineering quality, must grapple with structural limits that no amount of fine-tuning can fully resolve. The AI script writing app surfaced by Unsloppy is a concrete instance of the broader category: a tool designed to produce structured long-form narrative, operating within an architecture that has no native mechanism for tracking propositional commitment across scenes. The tool may produce fluent slug lines, plausible dialogue, and correctly formatted action description, but the underlying question is whether the output maintains the chain of referential obligations a working screenplay requires. This is not a complaint about the tool. It is a description of the architectural ceiling any such tool hits.

Consider a concrete scenario. A model is prompted to generate a twenty-page screenplay about a heist. In scene two, a character named Marcus hides a key under a flowerpot on the porch. By scene nine, the team needs to enter the house, and the model generates a scene where another character picks the lock. The key under the flowerpot — a prop explicitly established and therefore implicitly promised — has been forgotten. A human writer would either have the character use the key, or would have the flowerpot be empty, creating a new narrative beat, or would have the character not know about the key, creating dramatic tension. The model simply produces a lock-picking scene because, in the distribution of heist narratives in its training data, lock-picking scenes are plausible continuations. The model has not violated a commitment. It has no representation of commitment to violate.

Propositional Attitude Tracking and Architectural Absence

The deeper issue is that propositional attitudes — beliefs, desires, intentions, commitments — are not represented in transformer architectures in any structured way. The model’s attention mechanism can attend to earlier tokens, and in practice this means that with sufficient context length, a model can sometimes produce text that appears to reference earlier established facts. But this is not the same as maintaining a commitment state. Attention is a pattern-matching operation: it determines which earlier tokens are relevant to the current generation step based on learned associations. It does not encode the difference between ‘Vera warned about the door’ and ‘Vera is committed to the door being dangerous.’ These are different propositional attitudes toward the same content, and they have different consequences for what should happen next.

In a transformer, the representation of Vera’s warning is a vector in a high-dimensional space, contextualized by surrounding tokens. This vector can influence subsequent generation through attention, but it does not create a persistent state that constrains the model’s output space. There is no register that says ‘Vera is committed to P, so future outputs must be consistent with P or must explicitly address the inconsistency.’ The model can generate text that is inconsistent with P because nothing in the architecture prevents it. The warning is a pattern, not a promise.

This is related to what computational linguists call the commitment problem, though it goes by different names in different subfields. In discourse semantics, it is the problem of tracking what speakers are committed to across a dialogue or narrative. In formal semantics, it is the problem of representing propositional attitudes and their dynamics. In narrative theory, it is the problem of maintaining what literary scholar Marina Grishakova calls the ‘factuality gradient’ — the distinction between what is established as true in the story world, what is suspected, what is desired, and what is merely possible. Transformers collapse this gradient into a single representation: the token sequence. Everything that has been mentioned is, from the model’s perspective, equally part of the context. The door Vera warned about and the telephone that rang in scene five occupy the same representational status, even though one carries a narrative obligation and the other does not.

Why Coherence Metrics Miss the Failure

Standard evaluation frameworks treat screenplay generation as a coherence task. The assumption is that if a model produces text that is locally coherent — each sentence follows plausibly from the previous one — and globally coherent — the overall narrative has a recognizable structure — then the screenplay is successful. This assumption is wrong for the same reason that treating conversation as a coherence task is wrong: it confuses surface cohesion with referential accountability.

Cohesion is the surface-level linking of text — pronouns that refer to antecedents, conjunctions that signal logical relationships, lexical chains that maintain thematic continuity. Coherence, in the stronger sense, requires that the text be accountable to the commitments it has established. A screenplay can be highly cohesive — every sentence flows into the next, every scene heading is correctly formatted, every character name is consistent — and still fail as a screenplay because it does not honor the promises it has made. The door that was locked in scene three is opened without a key in scene twelve. The character who was in Chicago in scene seven is suddenly in Los Angeles in scene nine with no transition. The gun that was established as unloaded in scene two is fired in scene eleven without anyone reloading it.

These failures are invisible to n-gram overlap metrics because the offending scenes may share no n-grams with the establishing scenes. They are invisible to embedding similarity metrics because the scenes may be semantically related — both involve the same character, the same location, the same prop — even though the specific propositional content contradicts an earlier commitment. They are partially visible to human evaluation, but only if the evaluator is specifically instructed to track referential dependencies across the full document, which most evaluation protocols do not do.

The Authors Guild’s guidance on AI best practices for authors identifies a related concern from the professional writing community’s perspective: AI outputs are characterized as ‘generic mashups of pre-existing works’ that lack the voice, intentionality, and propositional commitment human-authored writing carries. The Guild’s framing implicitly identifies the same gap this article traces architecturally — the absence of speaker commitment, the difference between producing text that resembles a warning and actually committing to the proposition being warned about. Professional writers recognize this distinction at an industry level, even without the vocabulary of speech act theory. The difference between a screenplay that is accountable to its own promises and one that merely resembles such a screenplay is the difference between a document a production team can use and one they cannot.

The Illusion of Context Window Solutions

A natural response to the commitment problem is to note that modern models have context windows large enough to contain an entire screenplay. A 128,000-token context window can hold a 120-page script with room to spare. If the model can attend to every earlier token, should it not be able to maintain referential dependencies?

The answer is that attention is not commitment. A larger context window means the model can attend to the token sequence where Vera warned about the door, but it does not mean the model represents that warning as an active constraint. The model’s attention mechanism will distribute probability mass across relevant earlier tokens when generating scene twelve, but the relevance is determined by learned associations, not by a structured representation of what the earlier scene committed to. In practice, this means that models with large context windows still produce continuity errors in long-form generation, just at a lower frequency. The errors are less common because the relevant tokens are in the context, but they are not eliminated because there is no mechanism that enforces consistency as a constraint rather than as a statistical tendency.

This is why the problem is architectural, not parametric. Scaling up the model, expanding the context window, and increasing the training data all reduce the frequency of commitment violations, but they do not introduce a representation of commitment. The model becomes better at producing text that resembles a consistent screenplay, but it does not become a system that tracks what it has committed to. The distinction matters because it determines what kinds of errors are possible. A system that tracks commitments can fail to honor them, but its failures are systematic — it breaks specific obligations in specific ways. A system that does not track commitments fails differently: it produces text that is locally plausible but globally unaccountable, and the failures are not systematic because there is no system to violate.

What a Commitment-Aware Architecture Would Require

Specifying what is missing is easier than specifying what would fix it, but the diagnostic is valuable precisely because it clarifies the gap. A model that maintained propositional commitment would need, at minimum, a representation of the story world as a set of evolving states: what is true, what is believed by whom, what has been promised, what is physically present in each location. It would need a mechanism for updating this representation as new scenes are generated, marking certain facts as established, others as suspected, others as desired. And it would need a constraint mechanism that prevents the generation of scenes that violate established commitments without explicitly addressing the violation.

This is, in essence, a model of discourse state — what discourse semantics calls a commitment slate, adapted to the narrative domain. It is not a foreign concept in computational linguistics. Dialogue systems have long used representations of dialogue state to track what has been said and what obligations have been created. The difference is that screenplay generation requires a much richer state representation: not just what was said, but what the story world contains, what each character believes, what physical objects exist and where, what causal relationships have been established. This is closer to what cognitive science calls a situation model — a mental representation of the state of affairs described by a text — than to anything currently implemented in transformer architectures.

Some research directions point toward this kind of representation. Retrieval-augmented generation can maintain an external memory of established facts, though current implementations do not distinguish between committed propositions and mere mentions. Graph-based narrative generation systems maintain structured representations of plot state, but they typically operate at a higher level of abstraction than scene-level generation requires. Neurosymbolic approaches that combine neural generation with logical constraint satisfaction could, in principle, enforce commitment consistency, though no current system implements the full range of propositional attitudes screenplay generation demands.

The Question That Remains

The screenplay diagnostic reveals something shorter-form evaluation tends to obscure: language models produce text that carries the surface features of commitment without the underlying state. This is not a limitation scale alone will resolve, because it is not a problem of capacity. It is a problem of representation. The model can generate a warning, but it cannot be warned. It can generate a promise, but it cannot promise. It can generate a scene where a door is locked, but it cannot be committed to that door being locked.

Whether this matters depends on what the text is for. For many applications — drafting emails, generating marketing copy, producing boilerplate — the absence of commitment is invisible because the text is consumed locally and discarded. For screenplay generation, it is visible because the text must be accountable across a long structure, and the accountability is physically tested when the script goes into production. The diagnostic value of screenwriting is that it makes the gap between performance and understanding concrete: you can point to the door that was forgotten, the key that was misplaced, the vow that was ignored. These are not abstract philosophical failures. They are continuity errors a script supervisor would catch on the first read.

The question is whether the field will treat this as a problem worth solving or a limitation worth working around. Working around it means accepting that model-generated screenplays require human continuity checking, which is perhaps no different from the checking human-written screenplays require — except that human errors are systematic in ways that reveal the writer’s mental model of the story, while model errors are unsystematic in ways that reveal the absence of any mental model at all. Solving it means developing architectures that represent not just what was said but what was committed to, not just what tokens appeared but what obligations they created. That would require bridging the gap between locutionary production and illocutionary force at the architectural level, and it is not clear that any current approach can do so without a fundamentally different way of representing what language does. The door is still locked. The question is whether we remember why.