The Unit Everything Depends On
If there is a single concept that quietly determines what a large language model can and cannot do, it is the token. Tokens are rarely discussed outside technical circles, yet they shape nearly every strength, limitation, and odd behavior people encounter when using these systems. Without understanding tokens, it is easy to mistake fluent output for semantic understanding, or strange errors for malice or stupidity. In reality, many of these effects follow directly from what the model is actually manipulating.
A token is not a word, and it is not a unit of meaning. It is a computational fragment of text, chosen because it makes statistical learning efficient, not because it corresponds to how humans think or speak. Everything an LLM sees, processes, and produces is ultimately expressed as sequences of tokens.
From Text to Tokens
Before a model ever “reads” a sentence, that sentence is broken down into tokens using a process called tokenization. Depending on the language and tokenizer, a token might be a full word, a piece of a word, punctuation, whitespace, or even part of a number. For example, a single English word like “unbelievable” might be split into “un,” “believe,” and “able.” Rare or unfamiliar words are often broken into smaller fragments.
This is not a flaw; it is a design choice. Tokenization allows the model to handle an open-ended vocabulary without needing a separate symbol for every possible word. But it also means that the model never sees “words” in the way humans do. It sees symbol sequences, optimized for statistical reuse.
Crucially, tokenization is indifferent to meaning. Two tokens may look similar to a human but be unrelated statistically, while two fragments that feel unrelated to a reader may be closely linked in the model’s internal space because they often appear in similar contexts.
Why Tokens Are Not Semantic Units
Humans tend to assume that meaning attaches to words. Models do not make that assumption. Meaning, insofar as it exists for an LLM at all, is distributed across patterns of token co-occurrence. No single token “means” anything on its own. Its role is defined entirely by how it tends to appear next to other tokens.
This is why models can produce sentences that are locally coherent but globally nonsensical. Each step in generation may be statistically reasonable given the preceding tokens, even if the overall sentence contradicts itself or refers to impossible entities. The model is not tracking meaning across the sentence in the way a human reader would; it is tracking likelihoods across tokens.
This also explains why models struggle with tasks that require precise symbolic manipulation—counting characters, tracking variable names, or performing multi-step arithmetic reliably. Tokens are not symbols in the mathematical sense; they are fragments of language optimized for prediction, not calculation.
Tokens and the Illusion of Understanding
Because tokens often align approximately with words, and words carry meaning for humans, the output of an LLM can feel meaningful in a familiar way. This alignment is good enough to support conversation, explanation, and storytelling, but it is imperfect. The cracks appear in edge cases: unusual phrasing, rare names, long chains of reference, or tasks that require consistency across distant parts of a text.
From the model’s perspective, none of this is surprising. It is not failing to understand meaning; it was never operating on meaning in the first place. It is operating on statistical structure that usually correlates with meaning, until it doesn’t.
This distinction matters because it helps explain why adding more data or parameters improves fluency but does not produce comprehension. More data refines the statistics. It does not introduce a new representational layer where truth, intent, or reference suddenly appear.
Context Windows and Forgetfulness
Tokens also determine what the model can “remember.” Large language models operate within a context window, a finite number of tokens they can attend to at once. Anything outside that window is, from the model’s perspective, gone. There is no persistent memory unless it is explicitly simulated by external systems that re-inject prior text.
This leads to a common misconception: that models forget like humans forget. In reality, they do not forget; they simply cannot see beyond their current token window. When earlier details drop out of scope, the model continues as if they were never present. This can feel like inconsistency or deception, but it is neither—it is a direct consequence of token-bounded attention.
Understanding this also clarifies why “memory” features in AI products are not properties of the model itself, but of surrounding infrastructure that selects which tokens to reintroduce and when.
Tokens as a Design Constraint
The choice of tokenization scheme is not neutral. It affects how efficiently a model learns, how it handles different languages, how it represents numbers and code, and where it fails. Languages with complex morphology or non-alphabetic scripts can be disadvantaged or advantaged depending on how tokens are defined. Numeric reasoning often degrades because numbers are split into fragments that obscure their magnitude.
These effects are rarely visible to end users, but they shape outcomes in subtle ways. What looks like a reasoning failure is often a representational mismatch between human concepts and token-level structure.
Why This Matters
Tokens reveal a central truth about large language models: they do not operate on the units humans care about. They operate on units that make large-scale statistical learning possible. That gap is bridged by correlation, not understanding.
Once you see this, many debates become easier to resolve. The model is not secretly reasoning in a hidden semantic space. It is not manipulating ideas and then translating them into language. It is manipulating language fragments directly, and meaning emerges only insofar as language already encodes it.
In the next essay, we will turn to the process that fixes these token-level patterns into a usable system: training versus inference. Understanding when learning happens—and when it does not—will further dismantle the illusion that these systems are adapting, reasoning beings rather than highly capable statistical machines.
Member discussion: