← Back to concepts
7 min read

Tokens in large language models

Large language models do not read text as words or sentences. Everything a model processes — the prompt you send, the documents it retrieves, the reply it generates — is first broken down into tokens. The model receives token IDs, turns them into embedding vectors, and computes on those vectors.

What is a token?

A token is a chunk of text produced by a tokenizer: sometimes a whole word, sometimes part of a word, sometimes just punctuation or a single space. There’s no fixed rule like “one token per word” — the split depends on the tokenizer’s vocabulary, which was built from its training data.

"I love AI engineering." -> ["I", " love", " AI", " engineering", "."]   (5 tokens)
"unbelievable" -> ["un", "believ", "able"]                                (3 tokens)
"GPT-4" -> ["G", "PT", "-", "4"]                                          (4 tokens)

Tokenizers may also add special tokens that mark roles, tool calls, images, or the end of a message. Those markers are not visible words in your prompt, but they still count.

How text is split into tokens Three input strings enter the tokenizer. The tokenizer fans out illustrative splits: AI engineering into two common fragments, unbelievable into three subword fragments, and GPT-4 into four smaller pieces. The diagram shows that tokens can be whole words, word parts, punctuation, or spaces, and that exact splits vary by tokenizer. INPUT TEXT "AI engineering" "unbelievable" "GPT-4" Tokenizer looks up fragments in its vocabulary TOKENS "AI" " engineering" 2 tokens "un" "believ" "able" 3 subword tokens "G" "PT" "-" "4" Illustrative splits: exact token boundaries vary by tokenizer.
Tokens are the model's working chunks; these splits are illustrative because each tokenizer can split text differently.

Why not just count words or characters?

Words and characters are intuitive units for people, but they don’t match what the model actually processes:

  • Counting words undercounts, since many words split into two or more tokens (tokenization → token + ization).
  • Counting characters overcounts, since common fragments are compressed into single tokens (the, ing, tion are often one token each).
  • Counting tokens is the only number that matches the model’s real internal cost, because tokens are what the model’s context window, compute, and pricing are all measured in.

For example, the sentence “Please summarize this document for me” might look like 6 words, but tokenizes to something closer to 7-8 tokens once punctuation and word-fragments are accounted for.

Why tokens matter in practice

Tokens are the currency of everything an LLM does:

  • Context window: every model has a maximum number of tokens it can hold at once, covering the system prompt, conversation history, retrieved documents, and the model’s own output combined. A “128k context window” means 128,000 tokens total, not characters or words.
  • Latency: inference generation happens one token at a time, so longer outputs — more tokens — take longer to produce.
  • Cost: most API providers price per token, usually per 1,000 or per 1 million tokens, and often charge input and output tokens at different rates.
  • Prompt structure: if your system prompt, examples, and retrieved context together use most of the window, there’s less room left for conversation history or a long answer.

Consider a chatbot with a 16k-token context window. If the system prompt and retrieved documents already use 12,000 tokens, only about 4,000 tokens remain for conversation history and the model’s response — which might be just a few paragraphs.

How inputs and output share one token budget An illustrative 16000-token context window is shown as one shared bar. The system prompt uses 1500 tokens, retrieved documents use 10500 tokens, conversation history uses 1000 tokens, and only 3000 tokens remain for the model response. A second curated example keeps the system prompt, a smaller retrieval set, and a history summary, leaving a larger response reserve. Illustrative 16k-token context window limit Crowded request system 1.5k retrieved documents 10.5k chat 1k answer 3k Most of the window is already spent before the model writes. Curated request system 1.5k relevant chunks 5k summary 2k reserved for answer 7.5k Fewer input tokens leave more room for history and a complete response. Input tokens and output tokens compete for the same window.
A context limit is not just an input limit; it also has to leave enough tokens for the model's reply.

What happens when you exceed the window

A context window is a hard limit on the total tokens the model can process for one request. If the request plus the allowed output is too large, one of two things usually happens:

  • The API rejects the request with a context-length error.
  • The application silently truncates, summarizes, or drops older history before sending the request.

Many APIs also have a separate max output token setting. That setting caps how long the response may be, but it does not create extra room. A model with a 128k-token context window and a 4k output cap still needs the prompt, retrieved documents, history, hidden template tokens, and response budget to fit inside the total window.

This is why token budgeting is part of product design. You often need to reserve space for the answer, not just squeeze in as many documents as possible.

A worked example

Say you’re building a summarization feature and want to estimate cost before shipping it:

  • Input document: ~2,000 words ≈ 2,600 tokens (using the ~4 chars/token rule)
  • System prompt and instructions: ~150 tokens
  • Expected summary output: ~200 tokens

Total context usage is roughly 2,950 tokens. Billing separates that into 2,750 input tokens and 200 output tokens. If the model charges $3 per million input tokens and $15 per million output tokens, that single summarization costs:

input:  2750 / 1000000 x $3  = $0.00825
output:  200 / 1000000 x $15 = $0.00300
total:  $0.01125 (about $0.011)

Multiply by expected daily volume, and token counting turns into a real budgeting exercise, not just a technical curiosity.

Separating input and output token costs A summarization request has 2750 input tokens, made from a 2600-token document and 150 instruction tokens, and 200 output tokens. With illustrative prices of 3 dollars per million input tokens and 15 dollars per million output tokens, the input cost is 0.00825 dollars, the output cost is 0.003 dollars, and the total is about 0.011 dollars. Illustrative prices and token counts Input tokens 2750 2600 doc + 150 instr Input rate $3 / 1M Input cost $0.00825 2750 x $3 / 1M Output tokens 200 summary Output rate $15 / 1M Output cost $0.00300 200 x $15 / 1M Total about $0.011 Compute input and output costs separately, then add them.
Budgeting has to count input and output separately because providers often price them at different rates.

Try it yourself

The fastest way to build intuition is to tokenize real examples and see the exact split and count.

Try the OpenAI Tokenizer →

Paste in a paragraph you actually plan to send to a model — a support ticket, a code snippet, a prompt template — and compare the reported token count to your rough word-count estimate. Try the same text in a language other than English, too; many tokenizers use noticeably more tokens per word for non-English text since their vocabularies are trained mostly on English data.

Token efficiency varies by language and tokenizer. Languages under-represented in the tokenizer’s training data often need more tokens per word, but the only reliable answer is to measure with the exact tokenizer your model uses.

Measure real usage

Estimates help during design, but production systems should read the usage fields returned by the model API. These fields usually report input tokens, output tokens, and total tokens for the actual request.

Watch the provider’s definitions closely:

  • Cached input tokens may be counted or priced separately.
  • Some reasoning models count hidden reasoning tokens differently from visible output tokens.
  • Multimodal inputs such as images or audio also consume tokens, even when the user did not type text.
  • Chat templates and special tokens can add overhead that rough word-count estimates miss.

When designing an LLM feature, estimate token usage for realistic inputs, then measure real API usage and reserve enough context for the model’s response. Token counts are an engineering constraint, not just an implementation detail.