When I read a sentence, the words flow together as if they were always meant to be that way â each one a tidy packet of meaning. But the moment a machine touches that same sentence, it gets carved up. Sometimes into whole words, sometimes into slivers like âun-happi-ness,â and occasionally down to single letters. This carving-up, what the textbooks call tokenization, isnât just housekeeping. Itâs a foundational choice that sets the vocabulary, colors how morphology is understood, and draws the boundary of what can be learned. I keep returning to one question: what actually happens to the idea of âunhappinessâ when it lands as three pieces instead of one? The answer tells me as much about the limits of computation as it does about the stubborn nature of language.
”

The Origins of Tokenization: From Characters to Subwords
Early statistical language systems didnât bother with words. They just saw a parade of characters. An âaâ here, a âzâ there â each one a separate signal. There was a certain spare beauty to it. No need to know where word boundaries fell, no stumbling over misspellings or made-up terms. But that elegance came at a steep price. âThe cat sat on the matâ has 15 characters but only 6 words. A character-level model had to tiptoe across a dozen steps to catch a pattern that a word-level system could swallow in one go. Long-range dependencies turned into long, exhausting treks.
The obvious fix was word-level tokenization: keep a dictionary of known words, map each to a number, and shrink the sequences. It felt more human. But out-of-vocabulary words hit like a brick wall. Train on news articles, then bump into âsmartphoneâ or âblogosphere,â and all the model could do was slap an <UNK> token on them. Language never sits still â it coins, inflects, and mashes up terms without asking permission. A purely word-based vocabulary simply canât keep up.
That pressure gave us subword tokenization, the approach that now runs the show. Algorithms such as Byte-Pair Encoding (BPE) or SentencePiece donât commit to characters or words. They start with characters and, round after round, merge the most frequent pairs until a vocabulary of variable-length chunks emerges. Common words like âtheâ stay whole. Rare or twisty ones like âmisunderstandingsâ get split into âmis-under-stand-ings.â Nothing is ever completely unknown because anything can be dismantled into familiar subword units. Itâs a hybrid that feels, at times, like a clever compromise and, at others, like a bag of quirky accidents.
The Vocabulary as a Cognitive Filter
A vocabulary isnât a passive mirror. Itâs a filter that quietly decides which patterns will feel obvious to the model and which will feel like a struggle. Take âunlocking.â If the tokenizer keeps it whole, the model grabs its meaning as one compact embedding. If it becomes âun-lock-ing,â the model has to compose that meaning from three bits: the negation prefix, the root, and the progressive suffix. That compositional route can be a gift â it lets the model stretch to âunzippingâ or âunlockingâ even when those exact forms barely appeared during training. But it also demands that the model figure out the combinatorial rules of morphology, rules that are often leaky and context-dependent.
The vocabulary size itself sits there like a tuning knob with long-range effects. A tiny vocabulary â say, 8,000 tokens â forces heavy fragmentation. Each token carries less information, sequences grow long, and the model has to hold its attention across more steps to grasp meaning. A huge vocabulary â 50,000 tokens or more â leaves more words intact, shortens sequences, but fattens the embedding matrix and thins out the training signal across a long tail of rare tokens. Itâs a trade-off between compression and generalization, and it nudges the model toward its own signature mistakes.
What sticks with me is how tokenization can smuggle in linguistic bias without anyone noticing. A tokenizer fed mostly English text will merge âunâ and âlockâ into one token if that pair shows up often enough. The equivalent negation in a less-represented language may never get the same treatment. The model then reads that language with a clumsy, fragmented âaccent,â tripping over patterns that a better-fit vocabulary would have made plain. Same architecture, same training data â but the chopping method changes how the whole thing behaves.

Morphology, Meaning, and the Subword Lottery
Human languages build complex meanings with a kit of parts: prefixes, suffixes, roots, inflections. A subword tokenizer doesnât study morphology; it stumbles into it through statistics. The merges it learns follow frequency counts, not linguistic categories. In English, â-ingâ chases verb stems so relentlessly that it almost always ends up as its own token. The past tense â-edâ is messier. âWalkedâ becomes âwalk-ed,â a clean split. But âranâ? Thatâs a vowel change. No subword cut can isolate the pastness there. So the tokenizer leaves âranâ whole, handling two forms of the same grammatical feature in radically different ways.
This inconsistency has a cognitive echo. When a model reads âwalkedâ as two tokens, it might learn that â-edâ generally marks the past. When it reads âran,â that pattern is hidden â the past tense is baked into an atomic token. The model has to memorize âranâ as a separate dictionary entry, partly duplicating what it already knows about ârunningâ and âwalk.â Itâs not a bug so much as an inescapable constraint: the modelâs âsenseâ of morphology grows out of frequency accidents, not a grammar textbook.
Multilingual tokenization cranks up the volume on these effects. A vocabulary shared across dozens of languages has to ration its slots. English and Chinese grab a disproportionate share of whole-word tokens. Lower-resource languages get ground into smaller pieces. A Turkish sentence â where one word can bundle several agglutinated morphemes â becomes a long chain of subword tokens, each carrying only a sliver of meaning. That same concept in English might sit inside a single, dense token. The modelâs internal picture of the concept is then scattered across more vectors, harder to retrieve and work with.
Numerical Reasoning and the Character-Level Bottleneck
One of the stranger side effects of tokenization shows up with numbers. Numbers are infinitely productive and follow rigid compositional rules. A typical subword tokenizer, however, treats them like any other string. â237â might become a single token if it appears often enough. â238â might be split into â23-8â or â2-3-8,â depending on what the training data looked like. That inconsistency makes learning arithmetic a wobbly affair.
When â237â is its own token, the model can learn an embedding for it as a point in vector space. But it has no direct access to the fact that this number equals 200 + 30 + 7. Compare that to a digit-level representation, where â237â is always three tokens: â2,â â3,â â7.â There, the model can learn positional relationships â the â2â in the hundreds place shifts the value â but it has to build that understanding from scratch, without the crutch of a pre-merged token. Subword approaches land somewhere in the middle, sometimes grouping digits, sometimes not. The modelâs numerical reasoning gets brittle as a result.
That brittleness spills into other structured sequences: dates, times, codes, URLs. A tokenizer that splits â2024-11-17â one way in one corpus and another way in a different corpus forces the model to learn several surface forms for the same underlying thing. The inconsistency isnât random â it mirrors the statistics of the training data â but it stays opaque to the model, which has to treat each tokenization variant as a distinct input pattern.

Case, Punctuation, and the Boundaries of Sense
Tokenization also shapes how models perceive the surface of text: capitalization, punctuation, whitespace. Some tokenizers are case-sensitive, so âAppleâ and âappleâ sit in separate buckets with separate embeddings. Others lowercase everything, collapsing the two into one. The choice ripples out. A case-sensitive model can learn that âAppleâ the company and âappleâ the fruit are different things, but it might also grow jumpy about capitalization quirks it saw during training. A lowercased model loses that distinction but gets tougher against typos and casual writing.
Punctuation stirs up similar debates. A tokenizer that glues periods and commas to the preceding word â making âend.â a single token â treats the punctuation as part of the wordâs identity. One that splits punctuation off forces the model to learn its function as a structural marker. The first approach can blur the line between lexical meaning and syntactic boundaries; the second can make the model overly jumpy about missing or extra punctuation.
Whitespace handling may be the quietest influence of all. English leans on spaces to mark word edges; Chinese and Japanese do not. A tokenizer built with English in mind often assumes space-separated tokens and needs special handling for other scripts. Even inside English, contractions like âdonâtâ or âIâllâ raise questions. One token or two? The answer colors how the model learns negation, auxiliary verbs, and informal tone. Every small decision like this is a miniature theory about what counts as a meaningful unit of language.
Frequently Asked Questions
Why can’t we just use word-level tokenization with a giant vocabulary?
A vocabulary big enough to catch every possible word in a language would be unwieldy. New words appear all the time, and inflected forms multiply fast. Even if we assembled a vocabulary of millions of words, the model would struggle to learn useful representations for rare entries because they donât show up often enough during training. Subword tokenization gives a workable middle ground: it handles rare and novel words without falling apart, while keeping common words compact.
Does tokenization affect how well a model handles multiple languages?
Yes, and often unevenly. A multilingual tokenizer has to split its vocabulary budget across every language in the training data. Languages with more data tend to get more whole-word tokens; lower-resource languages get fragmented into smaller pieces. This means the model âreadsâ different languages at different granularities, which can affect fluency and the ability to carry knowledge from one language to another. The vocabulary size and the language mix of the training corpus are choices that matter a great deal.
Can a model learn to overcome bad tokenization?
To some extent, yes. Models can learn to attend across several tokens to rebuild a wordâs meaning, a bit like how we can read scrambled words. But bad tokenization makes learning clumsier. It stretches sequence length, hides regular patterns, and forces the model to spend capacity patching over the tokenizerâs oddities instead of picking up on higher-level relationships. A well-matched tokenizer doesnât guarantee success, but a poorly matched one is a persistent drag.
Is there any way to make tokenization completely language-agnostic?
Character-level tokenization comes closest, treating every script the same way. The catch is it creates very long sequences and shoves the whole job of learning word structure onto the model. Some researchers have been exploring byte-level models that work on raw UTF-8 bytes, doing away with a preset vocabulary entirely. Those look promising but currently demand more computation to reach similar performance. The hunt for a truly universal tokenization scheme is still very much underway.