What Do We Both Know? Stalnaker’s Common Ground as a Test for Dialogue Models

Stalnaker’s common ground is not a score. It is a state: the set of propositions that participants in a conversation take to be shared, and that they treat as available for presupposition and reference. A dialogue model can produce a fluent next turn without updating that state correctly. The useful question is therefore not whether a model “understands” common ground, but whether a controlled probe can show that it tracks a specific update.

This article describes a minimal-pair procedure for testing common-ground-sensitive behavior in dialogue models. The procedure separates three things that are often conflated in evaluation: factual recall, presupposition accommodation, and pragmatic inference. It also records the parameters needed to make a run reproducible.

Why common ground is a distinct target

Stalnaker’s account treats conversation as an activity that continually narrows the set of possible worlds compatible with what the participants accept as shared. A presupposition is a constraint on that set. When a speaker says my sister, the definite description presupposes that the addressee can identify a unique sister. If the addressee cannot, the utterance may still succeed through accommodation: the addressee adjusts the common ground to include the needed proposition.

Three failure modes follow, and they are not the same:

  • Factual failure. The model does not know the relevant fact (for example, that a named person has a sister).
  • Accommodation failure. The model knows the fact but does not treat the presupposition as something to be accommodated or challenged.
  • Pragmatic failure. The model accommodates, but does so in a way that violates the conversational purpose, for example by accepting a false presupposition that the addressee should reject.

A benchmark that only asks whether the reply is coherent will not distinguish these. A minimal-pair probe can.

A minimal-pair design

The design uses pairs of prompts that differ in one property of the common ground. The model’s output is compared across the pair. The comparison is the evidence; the absolute output is not.

Pair type 1: presupposition trigger present vs. absent

Prompt A: I ran into my sister at the market yesterday.
Prompt B: I ran into a friend at the market yesterday.

Prompt A contains a definite possessive that presupposes a unique sister. Prompt B does not. If the model’s next turn in A asks Which sister? or otherwise marks the presupposition, while its turn in B does not, the pair isolates sensitivity to the trigger. If both turns are identical, the model is not tracking the presupposition in this context.

Pair type 2: shared knowledge present vs. absent

Prompt A: We both know the meeting moved to Thursday. Can you confirm the room?
Prompt B: The meeting moved to Thursday. Can you confirm the room?

Prompt A explicitly marks the proposition as shared. Prompt B does not. A model that tracks common ground should treat the two differently: in A, the proposition is available for presupposition; in B, it may need to be introduced or confirmed. The probe records whether the model’s reply treats the proposition as already established.

Pair type 3: accommodation required vs. not required

Prompt A: My daughter’s violin recital is on Saturday. (The addressee has no prior information about a daughter.)
Prompt B: My daughter’s violin recital is on Saturday. (The addressee has just been told about the daughter.)

The two prompts are textually identical. The difference is in the prior context supplied to the model. If the model’s reply changes appropriately when the prior context establishes the daughter, the probe isolates accommodation from lexical sensitivity.

What to record

A common-ground claim is only as good as its run record. The following fields are the minimum for a reproducible probe. They follow the parameter names used in the Hugging Face GenerationConfig documentation, which is the reference for the generation parameters in the Transformers library.

Field Example Why it matters
model identifier meta-llama/Llama-3.1-8B-Instruct Different checkpoints behave differently.
revision commit hash or tag A model name alone does not pin the weights.
seed 42 Sampling is stochastic unless do_sample=False.
do_sample False Greedy decoding removes sampling variance.
temperature 1.0 Only relevant when sampling.
top_p 1.0 Nucleus sampling changes the output distribution.
top_k 50 Default in many configurations; record the actual value.
max_new_tokens 128 Truncation can look like a pragmatic failure.
prompt template exact string Small formatting changes alter behavior.
number of samples 20 Needed for any claim about a rate.

The Transformers documentation lists do_sample, temperature, top_k, top_p, max_new_tokens, and seed as generation parameters. Recording them is not optional if the claim is about model behavior rather than about a single output.

What the probe can and cannot show

A minimal-pair probe can show that a model’s output changes when a common-ground property changes, and that the change is in the expected direction. It cannot show that the model represents common ground as a set of propositions. That would require a different kind of evidence, for example an intervention that changes only the shared status of a proposition and observes a corresponding change in a downstream task.

The probe is also underpowered for small effects. If the pair differs by one word and the model’s output is identical in 19 of 20 samples, the probe has not shown insensitivity; it has shown that this pair, at this sample size, does not detect a difference. Report the interval, not just the point estimate.

Evaluation-validity audit

Many dialogue metrics reward surface continuity. A reply that repeats the user’s phrasing and stays on topic will score well even if it treats a non-shared proposition as shared. This is an evaluation-validity problem: the metric measures something other than the construct it is named for.

Two checks help:

  1. Construct check. Does the metric change when the common-ground status of a proposition changes, holding the surface form constant? If not, the metric is not tracking common ground.
  2. Adversarial check. Does the metric reward a reply that accommodates a false presupposition? If yes, the metric is not tracking successful communication.

These checks are cheap and can be run on any candidate metric. They do not require a new benchmark.

Relation to existing work

Chen et al. (2024) show that multi-step reasoning in several GPT-3 and GPT-4 variants fails two kinds of self-consistency: hypothetical consistency and compositional consistency. The finding is relevant here because common-ground tracking is a multi-step task. A model must represent what is shared, update it with each turn, and use the updated state to interpret the next utterance. If a model is inconsistent about its own hypothetical outputs, it is unlikely to maintain a stable common ground across turns.

Lin et al. (2022) introduce TruthfulQA, a benchmark of 817 questions across 38 categories designed to measure whether a model avoids generating false answers learned from imitating human texts. The best model in their evaluation was truthful on 58% of questions, while human performance was 94%. The relevance to common ground is indirect but important: a model that imitates popular misconceptions will also accommodate false presuppositions that reflect those misconceptions. Truthfulness and common-ground tracking are separate constructs, but they interact.

Feng et al. (2023) show that pretrained language models have political leanings that reinforce polarization in pretraining corpora, and that these leanings propagate into downstream hate speech and misinformation detectors. For common-ground probes, this is a warning about prompt selection. If the probe uses politically loaded examples, the model’s behavior may reflect a learned leaning rather than a common-ground computation. Use neutral examples for the core probe, and treat loaded examples as a separate condition.

Yao et al. (2020) train a contextual action language model on human gameplay and evaluate it on the Jericho benchmark, obtaining a 69% relative improvement in average game score over the previous state-of-the-art model. The relevance is methodological: their action candidates are conditioned on game history, which is a form of shared context. The lesson for common-ground probes is that the history supplied to the model must be controlled, because the model’s behavior depends on it.

A runnable probe skeleton

The following Python snippet uses the Transformers library. It is a skeleton, not a complete experiment. It records the parameters listed above and runs a pair of prompts.

from transformers import AutoModelForCausalLM, AutoTokenizer, GenerationConfig
import torch

model_id = "meta-llama/Llama-3.1-8B-Instruct"
revision = "main"  # replace with a commit hash for reproducibility
seed = 42

tokenizer = AutoTokenizer.from_pretrained(model_id, revision=revision)
model = AutoModelForCausalLM.from_pretrained(model_id, revision=revision, torch_dtype=torch.bfloat16)

gen_config = GenerationConfig(
    do_sample=False,
    max_new_tokens=128,
    temperature=1.0,
    top_p=1.0,
    top_k=50,
)

torch.manual_seed(seed)

prompts = [
    "I ran into my sister at the market yesterday.",
    "I ran into a friend at the market yesterday.",
]

for prompt in prompts:
    inputs = tokenizer(prompt, return_tensors="pt")
    outputs = model.generate(**inputs, generation_config=gen_config)
    print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Run this with do_sample=False first. If the outputs are identical across the pair, the probe has found no sensitivity under greedy decoding. Then run with sampling and 20 seeds to estimate a rate. Record both.

FAQ

Is common ground the same as shared attention?

No. Shared attention is about what participants are attending to. Common ground is about what they take to be shared. A model can attend to the same tokens as a human and still fail to treat a proposition as shared.

Can a single prompt test common ground?

No. A single prompt can show that a model produced a particular output. It cannot show that the output depends on common-ground status, because there is no comparison condition. Use a pair.

What if the model’s output is identical across the pair?

That is a result. It means the probe did not detect a difference at this sample size. Report it as such, and increase the sample size or change the pair before drawing a stronger conclusion.

Do I need a benchmark to test common ground?

No. A minimal pair with recorded parameters is sufficient for a narrow claim. A benchmark is useful when the claim is about a population of contexts, not a single pair.

What to do next

Pick one pair type. Run it with greedy decoding. If the outputs differ, run it with sampling and 20 seeds. Record the parameters. If the outputs do not differ, change one property of the pair and run again. The goal is not to prove that a model has common ground. The goal is to build a record of which common-ground updates a model tracks, under which parameters, and where it fails.

That record is more useful than a single score. It is also the only kind of evidence that can come out the other way.