Large language models cannot read text directly. Every prompt, document, or line of code has to be turned into numbers first, and that conversion is handled by a tokenizer.
What is a tokenizer?
A tokenizer is a program that splits raw text into small chunks called tokens, then looks each token up in a vocabulary — a fixed table that maps every known token to a unique integer ID.
"AI engineering" -> ["AI", " engineer", "ing"] -> [1723, 8842, 291]
The tokenizer passes those integer IDs to the model. The model then looks up an embedding vector for each ID, adds position information, and computes on those vectors. It never sees letters or words directly. Whatever the tokenizer decides to do with the input shapes everything the model computes afterward.
Here are a few more examples of how text commonly gets split:
"unbelievable" -> ["un", "believ", "able"]
"ChatGPT" -> ["Chat", "G", "PT"]
"2024" -> ["202", "4"]
"don't" -> ["don", "'t"]
"hello world" -> ["hello", " world"]
Notice that common English words like hello often stay whole, while brand names, numbers, and contractions get split in ways that can look unpredictable at first. That’s a direct result of which fragments were common enough in the training corpus to earn their own token.
Why not just split on words or letters?
A tokenizer needs a strategy for cutting text into pieces, and there are three broad options:
- Word-level: one token per word. Simple, but the vocabulary explodes with every new name, typo, or made-up word, and anything missing from the vocabulary can’t be represented.
- Character-level: one token per letter. Nothing is ever missing from the vocabulary, but ordinary sentences turn into very long sequences, which wastes context and compute.
- Subword: a middle ground. Most modern LLM tokenizers use subword methods such as BPE, often at the byte level.
Subword tokenization builds a vocabulary of frequently occurring fragments from a large body of training text. Common whole words like the or model usually get their own token, while rarer or longer words like tokenization get broken into reusable fragments such as token and ization. This lets the tokenizer represent essentially any input — including misspellings and words it has never seen — while keeping the vocabulary at a manageable, fixed size (often in the tens of thousands of entries).
For example, compare how a subword tokenizer might treat a familiar word versus an invented one:
"engineering" -> ["engineering"] (1 token - common enough to earn its own entry)
"engineeringhq" -> ["engineering", "hq"] (2 tokens - a new combination, split into known pieces)
"xqzzytron" -> ["x", "q", "zzy", "tron"] (4 tokens - nothing like it appeared in training)
The made-up word xqzzytron still gets represented, just less efficiently, falling back toward smaller and smaller fragments the less familiar it looks.
How the vocabulary gets built
Most subword tokenizers are trained before the model itself. The training process scans a huge text corpus and builds larger tokens from smaller pieces.
Byte-Pair Encoding (BPE) starts with small pieces, often bytes or characters, then repeatedly merges the most frequent adjacent pair. If l and o often appear together, the tokenizer may add lo; if lo and w often appear together, it may add low.
WordPiece is similar, but it scores merge candidates by how useful the combined piece is compared with the separate pieces. Unigram starts with many candidate pieces, then removes pieces that are least useful while keeping likely segmentations.
The result is a vocabulary tuned to the statistics of that corpus: common patterns in the training data become single tokens, and less familiar text falls back to smaller pieces.
This also means tokenizers trained on different data behave differently. A tokenizer trained mostly on English text will often split non-English words, code syntax, or rare technical jargon into more pieces than a tokenizer built with that content in mind. This is also why code often looks “expensive” to tokenize: symbols like {, }, and repeated indentation often consume separate tokens, though some tokenizers merge common whitespace runs. A short snippet of code can use noticeably more tokens than a sentence of plain English with the same character count.
Special tokens and byte fallback
Tokenizers also reserve IDs for special tokens. These tokens are not ordinary words from your prompt. They mark structure that the model was trained to understand:
- Beginning or end of a sequence.
- Chat role markers such as system, user, and assistant.
- Padding tokens used to make a batch the same length.
- Tool-call, tool-result, image, or audio placeholders in models that support them.
For example, a chat API may turn a simple conversation into a template like this before the model sees it:
Illustrative chat template:
<|system|>
You are a concise assistant.
<|user|>
Summarize this ticket.
<|assistant|>
Those markers are tokens too, and they count toward the context window. They also help the model separate instructions, user text, and its own reply during inference.
Modern byte-level tokenizers rarely need an unknown token. If a word, emoji, filename, or typo is not in the vocabulary as a whole piece, the tokenizer can fall back to smaller byte-level pieces that still reconstruct the original text. The tradeoff is efficiency: unfamiliar text may take more tokens.
Try it yourself
The best way to build intuition for tokenization is to see it happen on real text. OpenAI hosts a free interactive tool for exactly this — paste in a sentence and it will highlight each token in a different color and show you the token count and the underlying IDs.
A few things worth trying there:
- Type a common sentence and see how many words map one-to-one to tokens.
- Try a made-up word, a typo, or a name and watch it get split into smaller pieces.
- Paste a short block of code and compare its token count to a plain-English sentence of similar length.
- Try a non-English sentence and notice whether it uses more tokens than the English equivalent for the same idea.
Seeing these splits firsthand makes the rest of this concept path — tokens, context windows, and cost — much more concrete.
For production work, count tokens with the deployed model’s own tokenizer. Token counts differ between model families, and even closely related models may have different chat templates or special-token rules.
Why the tokenizer choice matters
The tokenizer sits at the very front of the pipeline, so its behavior ripples through everything downstream:
- Vocabulary size is a tradeoff: a larger vocabulary means fewer tokens per input but a bigger, more expensive embedding table for the model to store.
- Splitting behavior decides how many tokens a given piece of text becomes, which directly affects how much fits in the context window, how long inference takes, and API cost, since usage is billed per token. For example, the sentence “I love AI engineering” might cost 5 tokens, while a dense paragraph of unfamiliar jargon of similar length could easily cost double that.
- Tokenizer-model pairing is fixed: a model must always run with the exact tokenizer it was trained with. The vocabulary-to-ID mapping is baked in during training, so swapping tokenizers would make the model’s numeric IDs meaningless.
Common pitfalls
- Leading spaces matter: many tokenizers store
helloandhelloas different pieces. - Numbers can split oddly: dates, decimals, and IDs may break into chunks that do not match how people read them.
- Non-English text can cost more: token efficiency varies by language and tokenizer, especially for languages under-represented in the tokenizer’s training data.
- Letter counting is hard for LLMs: a model sees token pieces, not a clean character grid, so exact character tasks often need code or tools.
Once text has been tokenized, the model works with numeric tokens and the vectors derived from them — the next concept in this path.