Grammar and meaning are usually treated as separate layers of language. One is a system of formal rules; the other is a network of concepts and referents. In transformer language models, that separation gets harder to hold onto. A model that predicts the next token has to learn something about agreement, word order, and argument structure, but it also has to track what those forms point to. The question is not whether grammar and meaning interact in these systems. It is how tightly they are coupled, where the coupling breaks, and what those failures reveal about the kind of linguistic knowledge a network actually acquires. This article examines that relationship through cross-linguistic morphosyntactic evidence, evaluation design, and the limits of behavioral testing.

For readers of this blog, the stakes are concrete. If a model generalizes a subject-verb agreement pattern in English but fails on a parallel pattern in Basque or Swahili, the failure is not just a performance gap. It is evidence about whether the network has learned a general grammatical relation or a surface-level statistical regularity. The distinction matters for anyone using these models as tools for linguistic analysis, for typological comparison, or for building evaluation sets that claim to measure grammatical competence.
What Grammar and Meaning Look Like Inside a Transformer
Transformers do not store grammar as a set of rewrite rules. They encode token sequences as vectors, and each layer transforms those vectors through attention and feed-forward operations. A subject-verb agreement dependency, for example, is not represented as a labeled edge in a parse tree. It is distributed across attention patterns and hidden-state dimensions that change from layer to layer. When a model correctly produces the keys are on the table rather than the keys is on the table, the correct form emerges from a cascade of vector operations, not from a discrete agreement rule.
Meaning enters the same pipeline. A token like bank is disambiguated partly by nearby words, partly by learned co-occurrence statistics, and partly by whatever semantic structure the network has induced from its training data. The hidden state for bank after a few layers is not a clean bundle of grammatical features plus a clean bundle of semantic features. It is a single vector that blends both. That blending is why the grammar-meaning relationship in neural networks is an empirical question rather than a design assumption.
Evidence from Morphosyntactic Generalization
One way to probe the relationship is to test whether a model can generalize a morphosyntactic pattern to novel lexical items. In a controlled experiment, a network might be fine-tuned on sentences where a new verb takes a particular agreement suffix, then tested on sentences where that verb appears with a different subject. If the model produces the correct suffix, it has generalized the pattern. If it fails, the pattern may have been memorized as a lexical property rather than learned as a productive rule.
Cross-linguistic work complicates the picture. In languages with rich inflection, such as Finnish or Turkish, agreement and case marking carry more of the grammatical load than they do in English. A model that performs well on English subject-verb agreement may still fail on Finnish partitive case or Turkish vowel harmony. The failure pattern suggests that the network has not learned a universal notion of agreement. It has learned language-specific statistical regularities that happen to align with agreement in some languages and not in others.
This is not a claim that neural networks cannot learn morphology. They can, and they do. The claim is narrower: the kind of morphological knowledge a network acquires depends on the distributional properties of its training data, and those properties vary sharply across languages. A model trained mostly on English will not automatically transfer its English agreement behavior to a morphologically richer language, even when the underlying grammatical relation is similar.
Semantic Bootstrapping and Its Limits
There is a long tradition in linguistics of asking whether children use meaning to bootstrap grammar. A child who knows that dog refers to an animal and chase refers to an action can use that knowledge to infer that the dog chases the cat has a different structure from the cat chases the dog. The question for neural networks is whether they do something similar.
Some evidence suggests they do, but only partially. A transformer can use semantic cues to resolve syntactic ambiguity. In the chicken is ready to eat, the model’s interpretation depends on whether chicken is treated as the eater or the eaten. That decision is not purely syntactic. It draws on world knowledge about chickens, eating, and typical sentence roles. When the semantic cue is strong, the model often chooses the pragmatically plausible reading. When the cue is weak or contradictory, performance drops.
The limit is that semantic bootstrapping in a neural network is not the same as semantic bootstrapping in a child. A child uses meaning to build a grammar that then becomes productive. A transformer uses meaning as one more statistical signal in a prediction task. The grammar that emerges is not a separate system that has been bootstrapped. It is a byproduct of the same vector operations that handle meaning. That is why the two are so hard to disentangle in practice.
Evaluation Validity and the Grammar-Meaning Confound
Many evaluation sets for grammatical competence are built by creating minimal pairs. One sentence is grammatical; the other differs by a single word or morpheme. The model is scored on whether it assigns a higher probability to the grammatical sentence. The method is clean in principle. In practice, it often confounds grammar with meaning.
Consider a minimal pair like the cat is sleeping versus the cat are sleeping. The ungrammatical sentence is not just grammatically wrong. It is also statistically unusual, semantically odd, and pragmatically implausible. A model that prefers the grammatical sentence may be using any of those signals. The result tells us little about whether the model has learned subject-verb agreement as a grammatical relation.
Cross-linguistic minimal pairs make the confound worse. In a language with flexible word order, a minimal pair may differ in information structure rather than grammaticality. In a language with pro-drop, a missing subject may be grammatical in some contexts and ungrammatical in others. If the evaluation set does not control for these factors, the score reflects a mixture of grammatical, semantic, and discourse-level knowledge. The grammar-meaning relationship becomes a source of measurement error.
Probing as a Partial Solution
Probing classifiers are one attempt to separate the two. A probe is a small model trained on top of a transformer’s hidden states to predict a linguistic property, such as part of speech, dependency label, or semantic role. If the probe succeeds, the argument goes, the hidden states must contain information about that property. If the probe fails, the information is absent or not linearly separable.
Probing has limits. A probe can succeed because the hidden states contain the target property, or because the probe itself learns to extract a correlated property. A probe trained to predict subject-verb agreement might succeed by detecting number features on the subject noun, not by detecting the agreement relation itself. The distinction matters for the grammar-meaning question. If a probe succeeds by using semantic features, then the result does not show that grammar is represented independently of meaning. It shows the opposite.
Control tasks help. A common control is to train the probe on a randomized version of the labels. If the probe performs no better on the real labels than on the randomized ones, the original result is suspect. Another control is to train the probe on a related but distinct property. If a probe trained to predict agreement also predicts semantic number, the two properties are entangled in the hidden states. That entanglement is itself a finding about the grammar-meaning relationship.

Cross-Linguistic Failure Analysis
Failure analysis is more informative than aggregate accuracy. When a model fails on a particular construction, the pattern of failure tells us what it has learned. A model that fails on English subject-verb agreement only when the subject and verb are separated by a long relative clause is not failing on agreement per se. It is failing on long-distance dependency tracking. A model that fails on Finnish case marking only for low-frequency nouns is failing on lexical coverage, not on case as a grammatical category.
Cross-linguistic failure analysis adds another dimension. The same model may fail on subject-verb agreement in English, Basque, and Swahili, but for different reasons. In English, the failure may be due to intervening nouns. In Basque, it may be due to ergative case marking. In Swahili, it may be due to noun class prefixes. The failures are not three instances of one problem. They are three different problems that happen to share a label.
This is why the grammar-meaning relationship cannot be studied in a single language. A model’s behavior in English is shaped by English-specific distributional facts. A model’s behavior in a typologically different language is shaped by different facts. The relationship between grammar and meaning that emerges in one language may not generalize to another. The cross-linguistic evidence is the only way to see the full shape of the relationship.
Case Study: Agreement Attraction
Agreement attraction is a well-documented phenomenon in human sentence processing. In the key to the cabinets are on the table, the plural noun cabinets attracts the verb into plural agreement, even though the grammatical subject is singular. Humans make this error under time pressure. Transformers make it too, but the pattern differs across languages and across model sizes.
In English, larger models show less attraction than smaller models, but the effect does not disappear. In languages with richer agreement morphology, the attraction pattern interacts with case and gender features. A model that resists attraction in English may still be pulled by a plural dative noun in Russian. The cross-linguistic pattern suggests that attraction is not a single grammatical phenomenon. It is a family of interference effects that depend on how a language encodes number, case, and gender.
For the grammar-meaning question, attraction is a useful test case because it sits at the boundary between the two. The attractor noun is grammatically irrelevant but semantically salient. A model that is pulled by the attractor is letting a semantic or distributional signal override a grammatical one. The size of the attraction effect is a measure of how tightly grammar and meaning are coupled in the network’s representations.
Philosophy-of-Language Diagnostics
The grammar-meaning relationship in neural networks also raises questions that belong to philosophy of language. One is the question of compositionality. A compositional system builds the meaning of a sentence from the meanings of its parts and their syntactic combination. A transformer does not explicitly build meanings this way. It computes a sequence of vector transformations. Whether the result is compositional in any interesting sense is an empirical question, not a definitional one.
Another question is about the distinction between competence and performance. In linguistics, competence is the idealized knowledge of a language; performance is what a speaker actually does. A transformer has no competence in this sense. It has only performance, in the form of next-token predictions. Any grammar we attribute to it is an inference from that performance. The inference may be useful, but it is not the same as discovering a rule system inside the network.
A third question is about reference. A transformer that produces the cat is on the mat is not referring to a particular cat or a particular mat. It is producing a sequence of tokens that a human reader can interpret as referring. The network’s relationship to meaning is mediated by the human interpreter. That mediation is easy to forget when the model’s output looks fluent. The grammar-meaning relationship inside the network is not the same as the grammar-meaning relationship in a human speaker, because the network is not a speaker in the relevant sense.
What the Evidence Does and Does Not Show
The evidence supports a few cautious conclusions. First, grammar and meaning are entangled in transformer representations. They are not cleanly separable modules. Second, the degree of entanglement varies across languages and across construction types. Third, evaluation methods that assume a clean separation will overestimate grammatical competence when semantic cues are available and underestimate it when they are not. Fourth, cross-linguistic failure analysis is the most reliable way to see the relationship clearly, because it controls for language-specific confounds.
The evidence does not support stronger claims. It does not show that transformers have no grammatical knowledge. It does not show that they have human-like grammatical knowledge. It does not show that meaning is primary and grammar is secondary, or the reverse. The relationship is more like a gradient than a hierarchy. Some constructions are more grammar-driven; others are more meaning-driven. The network’s behavior reflects that gradient, not a fixed architecture.
Practical Takeaways for Evaluation and Analysis
For anyone building evaluation sets or analyzing model behavior, the grammar-meaning confound is a design problem with concrete solutions. One solution is to use cross-linguistic minimal pairs that control for semantic plausibility. A minimal pair in which both sentences are semantically plausible but only one is grammatical isolates grammar more cleanly than a pair in which the ungrammatical sentence is also semantically odd.
Another solution is to report failure patterns, not just accuracy. A model that scores 90% on an agreement test may be failing on exactly the constructions that matter for a theoretical claim. Reporting which constructions fail, and how the failures pattern across languages, is more informative than a single number.
A third solution is to use multiple probe types with control tasks. A probe that succeeds on agreement but also succeeds on a semantic control task is not evidence for a clean grammatical representation. The control result should be reported alongside the main result, not buried in an appendix.
These practices do not eliminate the grammar-meaning confound. They reduce it. The confound is not a bug in the method. It is a fact about the object of study. Grammar and meaning are entangled in neural networks because they are entangled in the distributional statistics those networks learn from. The goal is not to pretend the entanglement does not exist. The goal is to measure it accurately.

Frequently Asked Questions
Do transformers learn grammar separately from meaning?
No. The available evidence suggests that grammatical and semantic information are distributed across the same hidden-state vectors and attention patterns. A transformer does not have a grammar module and a meaning module. It has a single sequence of vector transformations that encodes both. Probes can sometimes separate the two statistically, but the separation is partial and depends on the probe design.
Why do cross-linguistic tests matter for the grammar-meaning question?
Cross-linguistic tests matter because a model’s behavior in one language is shaped by that language’s distributional properties. A pattern that looks like grammatical generalization in English may be a surface-level statistical regularity that does not transfer to a language with different morphology or word order. Testing across languages reveals whether a behavior is general or language-specific.
Can a model use meaning to bootstrap grammar the way children do?
Only in a limited sense. A transformer can use semantic cues to resolve syntactic ambiguity and to prefer plausible interpretations. But the grammar that emerges is not a separate productive system that has been bootstrapped from meaning. It is a byproduct of the same prediction task that handles meaning. The two are entangled from the start.
What is the biggest risk in evaluating grammatical competence?
The biggest risk is the grammar-meaning confound. Many evaluation sets use minimal pairs where the ungrammatical sentence is also semantically odd or statistically unusual. A model that prefers the grammatical sentence may be using semantic or distributional cues rather than grammatical knowledge. The result overestimates grammatical competence. Cross-linguistic controls and failure-pattern reporting reduce this risk.
Next Steps for This Line of Inquiry
The natural next step for this blog is a closer look at agreement attraction across languages. The phenomenon sits at the grammar-meaning boundary and has a rich experimental literature in human sentence processing. A follow-up article could compare attraction effects in English, Russian, and Swahili, using the same model and the same evaluation protocol. That would turn the general claims in this article into a concrete, testable case study.
Another path is a glossary entry on the grammar-meaning confound. The term appears throughout this article, but it deserves a standalone definition with examples from evaluation design. A glossary entry would give future articles a stable reference point and help readers who encounter the term in other contexts.
Both paths strengthen the same editorial thesis: that cross-linguistic failure analysis, not aggregate accuracy, is the most reliable way to understand what transformer language models actually learn about grammar and meaning.