How Language Models Handle Ambiguity and Why It Matters

Ambiguity is a property of natural language in which a single surface form can correspond to multiple underlying structures or meanings. In morphosyntactic terms, this includes lexical ambiguity, structural ambiguity, and scope ambiguity. For transformer language models, ambiguity is not a peripheral edge case; it is a central test of whether a model has acquired a grammar that generalizes beyond surface statistics. This article examines how transformer models assign probability to ambiguous strings, what that reveals about their internal representations, and why evaluation protocols must be designed with ambiguity in mind.

For researchers working on cross-linguistic failure analysis, ambiguity offers a controlled diagnostic. A model that assigns high probability to only one reading of a globally ambiguous sentence may be relying on shallow heuristics. A model that assigns balanced probability to multiple readings may be tracking genuine structural alternatives. The distinction matters for claims about syntactic competence, semantic composition, and the validity of benchmark scores.

What Ambiguity Tests Actually Measure

Ambiguity tests are not a single method. They range from human acceptability judgments to targeted probing of model probabilities. The most informative designs compare a model’s behavior on ambiguous strings with its behavior on minimally different unambiguous controls. For example, a sentence like The woman saw the man with the telescope has two attachment sites for the prepositional phrase. A model that assigns similar probability to both readings is not necessarily confused; it may be representing both parses. A model that assigns sharply higher probability to one reading may be showing a systematic bias.

In cross-linguistic work, ambiguity interacts with language-specific properties. Japanese allows argument ellipsis that English does not. German allows scrambling that changes surface order without changing coreference. A transformer trained on multilingual data may show different ambiguity resolution patterns in each language, even when the surface string is superficially similar. This makes ambiguity a useful probe for whether a model has acquired language-specific constraints or is applying a language-general fallback strategy.

Lexical Ambiguity and Frequency Effects

Lexical ambiguity is the simplest case: a word form like bank can refer to a financial institution or a river edge. Transformer models typically assign higher probability to the more frequent sense in a neutral context. This is not surprising, but it becomes interesting when context is added. A model that fails to shift probability toward the river sense after the word fishing is not integrating context the way a human parser would. The failure is not a lack of lexical knowledge; it is a failure of contextual update.

Frequency effects are not uniform across languages. In morphologically rich languages, a word form may be ambiguous between different case or agreement features. A model that resolves this ambiguity by frequency alone will make errors that a human speaker would not. This is a concrete, testable prediction that can be checked with existing multilingual corpora.

Structural Ambiguity and Attachment Preferences

Structural ambiguity arises when a string can be parsed into more than one syntactic tree. The classic example is prepositional phrase attachment. In English, the sentence I saw the man with the telescope can mean that the speaker used a telescope or that the man had a telescope. Human readers show a mild preference for one reading, but the preference is modulated by verb semantics, definiteness, and discourse context.

Transformer models show attachment preferences that are often correlated with surface linear order. A model may prefer to attach a prepositional phrase to the most recent noun phrase, regardless of verb semantics. This is a heuristic that works often but fails systematically. When the heuristic fails, the model’s probability distribution over readings becomes a diagnostic of its parsing strategy. A model that assigns near-zero probability to the dispreferred reading is not representing ambiguity; it is collapsing it.

Why Ambiguity Matters for Evaluation Validity

Most benchmark datasets for language models are built from naturally occurring text. This means they contain ambiguous strings, but the ambiguity is rarely annotated. A model can score well on a benchmark by consistently choosing the most frequent reading, even if it never represents the alternative. This inflates benchmark scores and obscures the model’s actual competence.

Evaluation validity requires that we know what a test item is testing. An ambiguous sentence in a benchmark is not a clean test of anything unless the ambiguity is controlled. If a dataset includes The woman saw the man with the telescope as a test item, the model’s answer depends on which reading the dataset creator intended. If the intended reading is not specified, the item is uninterpretable. This is a measurement problem, not a model problem.

Cross-linguistic evaluation adds another layer. A sentence that is ambiguous in English may be unambiguous in Japanese, or vice versa. A benchmark that translates test items without checking for ambiguity shifts is not measuring the same thing across languages. This is a known issue in multilingual evaluation, and it is one reason why cross-linguistic failure analysis must be done with carefully constructed minimal pairs.

What Transformer Representations Reveal

Transformer models do not parse sentences in the way a symbolic parser does. They assign probability to tokens based on contextualized representations. When a model processes an ambiguous string, the representation at each position is a blend of possible continuations. This blend can be probed by looking at the model’s next-token distribution or by using diagnostic classifiers.

One finding from recent work is that transformer models often retain information about multiple readings in their hidden states, even when their final output distribution is sharply peaked. This suggests that the model is not failing to represent ambiguity; it is failing to surface it. The distinction matters for interpretability. A model that represents both readings but outputs only one is different from a model that never represents the alternative.

This has implications for how we think about model competence. A model that represents both readings but outputs only one may be doing something like human garden-path recovery. A model that never represents the alternative is doing something more like heuristic pattern matching. The two cases require different interventions and different evaluation methods.

Cross-Linguistic Patterns in Ambiguity Resolution

Languages differ in which ambiguities they tolerate and how they resolve them. English tolerates a great deal of structural ambiguity because word order is relatively fixed. Japanese tolerates argument ellipsis because case markers and discourse context often disambiguate. German tolerates scrambling because case morphology marks grammatical roles. A transformer model trained on all three languages may show different resolution patterns in each, but the patterns are not always language-appropriate.

For example, a model may resolve Japanese ellipsis by relying on linear order, even though Japanese word order is less informative than case marking. This is a cross-linguistic failure that would not show up in English-only evaluation. It is also a failure that matters for downstream tasks like machine translation and cross-lingual transfer. If a model resolves ellipsis incorrectly in Japanese, it will produce incorrect translations into English, where the elided argument must be overtly realized.

These cross-linguistic patterns are not random. They tend to follow the model’s training data distribution. Languages with more training data show more human-like resolution patterns. Languages with less training data show more surface-order heuristics. This is a predictable consequence of statistical learning, but it is not a trivial one. It means that ambiguity resolution is a place where data imbalance becomes visible as a competence gap.

Practical Diagnostics for Ambiguity

Researchers who want to test how a model handles ambiguity can use a few practical diagnostics. The first is minimal pair comparison. Take an ambiguous sentence and create two unambiguous paraphrases, one for each reading. Compare the model’s probability for the ambiguous sentence with its probability for each paraphrase. A model that assigns similar probability to the ambiguous sentence and both paraphrases is representing both readings. A model that assigns high probability to only one paraphrase is collapsing the ambiguity.

The second diagnostic is continuation probability. Give the model an ambiguous sentence and ask it to continue the text. The continuation will often disambiguate the reading. If the model’s continuation is consistent with only one reading, that tells you which reading the model committed to. If the continuation is consistent with both readings, the model is maintaining ambiguity.

The third diagnostic is cross-linguistic comparison. Take an ambiguous sentence in one language and its translation in another language where the ambiguity is resolved. Compare the model’s behavior on both. If the model resolves the ambiguity differently in the two languages, that is evidence of language-specific processing. If it resolves them the same way, that is evidence of a language-general heuristic.

Limitations and Open Questions

The evidence on ambiguity in transformer models is partial. Most studies use English, and most use a small set of hand-constructed examples. The cross-linguistic picture is even thinner. We do not yet know whether the patterns found in English generalize to other languages, or whether they are an artifact of English-specific training data.

There is also a question about what counts as a correct resolution. Human speakers do not always agree on the preferred reading of an ambiguous sentence. Some ambiguities are genuinely unresolved in context. A model that assigns balanced probability to two readings may be doing exactly what a human would do. A model that assigns sharply peaked probability may be overcommitting. The evaluation standard must be calibrated to human behavior, not to a single gold reading.

Finally, there is the question of whether ambiguity resolution is a single ability or a collection of related abilities. Lexical ambiguity, structural ambiguity, and scope ambiguity may be handled by different mechanisms. A model that is good at one may be bad at another. Treating them as a single phenomenon risks obscuring the specific failure modes that matter for cross-linguistic generalization.

FAQ

What is the difference between lexical and structural ambiguity?

Lexical ambiguity occurs when a single word form has multiple meanings, such as bank meaning a financial institution or a river edge. Structural ambiguity occurs when a string of words can be parsed into more than one syntactic tree, such as I saw the man with the telescope, where the prepositional phrase can attach to the verb or the noun. Transformer models often resolve lexical ambiguity by frequency and structural ambiguity by linear order, but both strategies fail in predictable ways.

Why do benchmark scores not reflect ambiguity handling?

Benchmark datasets are typically built from naturally occurring text, which contains ambiguous strings that are not annotated for their intended reading. A model can score well by consistently choosing the most frequent reading, even if it never represents the alternative. This inflates benchmark scores and obscures the model’s actual competence. Evaluation validity requires controlled minimal pairs where the intended reading is specified.

How can researchers test ambiguity in transformer models?

Three practical diagnostics are useful. First, compare model probability for an ambiguous sentence with probability for unambiguous paraphrases of each reading. Second, examine continuation probability to see which reading the model commits to. Third, compare behavior on an ambiguous sentence in one language with its translation in a language where the ambiguity is resolved. These methods reveal whether the model represents multiple readings or collapses them.

Do transformer models represent ambiguity even when they output one reading?

Evidence from probing studies suggests that transformer models often retain information about multiple readings in their hidden states, even when their final output distribution is sharply peaked. This means the model may represent ambiguity but fail to surface it. The distinction matters for interpretability and for deciding whether a failure is a parsing error or an output-layer bias.

Next Steps for This Site

This article is part of a series on morphosyntactic generalization in transformer models. A natural follow-up is a cross-linguistic study of attachment preferences in English, German, and Japanese, using the minimal pair method described above. Another path is a glossary entry on scope ambiguity, which is closely related but involves quantifier scope rather than phrase attachment. Reader questions on specific ambiguity types are welcome and will be addressed in future posts.

A person reading a book with ambiguous text highlighted
A chalkboard with syntactic tree diagrams showing two possible parses
A researcher comparing language data on a computer screen