← Back to concepts
10 min read

Inference — how models generate output

Inference is what happens when you send input to a trained model and receive output back. Training changes model weights. Inference uses the weights that already exist.

For a language model, inference usually means taking a prompt, processing its tokens, and generating a response one token at a time. The model may be an LLM in a chat product, a code assistant, or a batch job that summarizes documents.

From prompt to prediction

Before the model can answer, the input is converted into tokens. Those tokens move through the model’s layers, where attention and learned weights produce a representation of the current context. That context can include instructions, user messages, retrieved documents, tool results, and earlier output.

At the end of that pass, the model does not directly produce a finished paragraph. It produces logits, raw scores for every token in its vocabulary. A softmax turns those scores into probabilities, which answer a question like:

Given everything so far, what token is most likely to come next?

If the prompt is:

The capital of France is

the model may assign a very high probability to Paris. Other tokens may still have some probability, but they are less likely. The exact distribution depends on the model architecture, weights, prompt, and decoding settings.

Generation is a loop

Text generation repeats the same basic loop:

  1. Read the current input tokens.
  2. Compute scores for the next token at the last position.
  3. Choose one token using a decoding strategy.
  4. Append that token to the context.
  5. Repeat until the model stops or reaches a limit.

For example, this illustrative trace shows token-by-token growth, not a full sentence written at once:

Prompt:  "Write one sentence about embeddings."
Step 1:  "Embeddings"
Step 2:  "Embeddings turn"
Step 3:  "Embeddings turn data"
Step 4:  "Embeddings turn data into"
Step 5:  "Embeddings turn data into vectors"

The model is not writing the whole sentence in one operation. It is repeatedly choosing the next token based on the prompt and everything it has already generated. Each decoding step produces scores for the next token at the last position, and one token is chosen from those scores.

The token-by-token generation loop The current context, ending in Embeddings turn, passes through the model's forward pass with fixed weights. The model outputs next-token probabilities, such as data 0.41, text 0.22, inputs 0.09, and all others 0.28. Decoding picks one token, data. A stop check asks whether a stop sequence, end token, or max token limit was reached. If not, the chosen token is appended to the context, where it can be streamed to the user, and the loop repeats. If yes, the finished response is returned. Current context "...Embeddings turn" Model forward pass, fixed weights Probabilities data 0.41 text 0.22 inputs 0.09 top 3; others 0.28 Decoding picks one token chosen: "data" Stop? stop sequence, end token, or max tokens reached no Append token to context "...Embeddings turn data" each new token can be streamed to the user repeat yes: finished response
Each decoding step scores the next token, chooses one token, appends it, and repeats until a stop condition ends generation.

Prefill, decode, and the KV cache

Autoregressive inference has two phases.

Prefill processes the prompt tokens. In most transformer models, the prompt tokens can be processed in parallel in one forward pass. This phase drives time to first token, the delay before streaming starts.

Decode generates output tokens. Each new token depends on the tokens before it, so output appears one token at a time. This phase drives tokens per second, the speed of the stream after it starts.

A KV cache makes decode faster. In attention layers, every token creates keys and values. The cache stores those keys and values so the model does not recompute them for earlier tokens at every step. This trades memory for speed. The cache grows with context length and batch size, so very long prompts and many concurrent requests can become memory-bound.

Prefill, decode, and a growing KV cacheThe prompt tokens are processed together during prefill, producing the first next-token scores and filling the KV cache with keys and values for prompt tokens. Decode step one chooses token A and adds its keys and values to the cache. Decode step two uses the existing cache plus token A to choose token B and appends another cache entry. The cache grows as context and output grow.PREFILL: PROCESS PROMPT IN PARALLELPrompt tokens[Write] [about] [vectors]Model passscores first tokenKV cache after prefillK,V for 3 prompt tokensDECODE: ONE OUTPUT TOKEN PER STEPStep 1use cache, choose AKV cache growsprompt K,V + token AStep 2use cache, choose Bappend new K,V each tokenPrefill affects time to first token. Decode speed affects tokens per second. Cache memory grows with context x batch.
The KV cache avoids recomputing earlier keys and values, which speeds decoding but uses more memory as the context grows.

Decoding controls the next token

The model produces probabilities, but the application still has to decide how to pick from them. This choice is called decoding.

Common decoding strategies include:

  • Greedy decoding: always pick the highest-probability token. It is simple and repeatable, but it can get stuck in bland or locally good choices.
  • Sampling: draw from the probability distribution. Temperature, top-k, and top-p shape which tokens are likely enough to sample.
  • Beam search: keep several candidate continuations and score them as they grow. It is common in translation and other encoder-decoder systems, but less common for open-ended chat because it can reduce diversity.

Common decoding settings include:

  • Temperature: divides the logits before softmax. Low temperature sharpens the distribution; high temperature flattens it.
  • Top-p: limits choices to the smallest group of tokens whose combined probability reaches a threshold.
  • Top-k: limits choices to the k most likely tokens.
  • Max output tokens: caps how long the response can be.
  • Stop sequences: tell the system when to stop generation after a specific pattern appears.
How top-k and top-p filter candidate tokens Five visible candidate tokens continue The server failed because, with illustrative normalized probabilities the 0.47, a 0.26, of 0.12, there 0.09, and database 0.06. Top-k with k equal to 2 keeps only the two most likely tokens, the and a. Top-p with p equal to 0.80 adds tokens in order until the running total reaches 0.80: 0.47, then 0.73, then 0.85, so it keeps the, a, and of. The remaining tokens are removed before one token is sampled. Candidates after "The server failed because" (five-token view, illustrative) Top-k = 2 keep the k most likely tokens TOKEN PROB RESULT "the" 0.47 kept "a" 0.26 kept "of" 0.12 removed "there" 0.09 removed "database" 0.06 removed Top-p = 0.80 keep the smallest set reaching p TOKEN PROB TOTAL RESULT "the" 0.47 0.47 kept "a" 0.26 0.73 kept "of" 0.12 0.85 kept "there" 0.09 0.94 removed "database" 0.06 1.00 removed Dashed line = cutoff. Kept probabilities are renormalized, then one token is sampled from them.
Top-k keeps a fixed number of candidates, while top-p keeps however many are needed to cover a probability budget.

For a customer support classification task, you usually want low randomness. For brainstorming names for a product, you may want more variety.

A worked example

Imagine a model has to continue this prompt:

The server failed because

It might assign top-candidate probabilities like this:

"the"                  0.32
"a"                    0.18
"of"                   0.08
"there"                0.06
"database"             0.04
all other tokens        0.32

Those five named candidates sum to 0.68 because the remaining probability is spread across many other tokens. The temperature diagram below uses the same five named candidates but renormalizes only that small visible set to 1.0, so the numbers are not the full vocabulary distribution.

Temperature divides logits before softmax. With a very low temperature, sampling will likely choose the because the distribution becomes sharp. With greedy decoding, the top token is always chosen regardless of temperature. With a higher temperature, sampling may choose another plausible token. After choosing one token, the model runs again with the longer context.

How temperature reshapes next-token probabilities The five named candidate tokens from the worked example are shown at three temperatures, renormalized over those five tokens. At temperature 0.5 the distribution is sharper: the 0.70, a 0.22, of 0.04, there 0.02, database 0.01. At temperature 1.0 it keeps the normalized five-token shape: the 0.47, a 0.26, of 0.12, there 0.09, database 0.06. At temperature 2.0 it is flatter: the 0.33, a 0.25, of 0.16, there 0.14, database 0.12. The ranking never changes, but lower temperature concentrates probability on the top token and higher temperature spreads it out. Temperature 0.5 sharper: strongly favors "the" "the" 0.70 "a" 0.22 "of" 0.04 "there" 0.02 "database" 0.01 Temperature 1.0 baseline: normalized five tokens "the" 0.47 "a" 0.26 "of" 0.12 "there" 0.09 "database" 0.06 Temperature 2.0 flatter: more variety "the" 0.33 "a" 0.25 "of" 0.16 "there" 0.14 "database" 0.12 Computed from the five named candidates only, renormalized and rounded. Same bar scale in every panel.
Temperature never reorders the candidates; it changes how strongly the top token dominates sampling.

This is why the same prompt can produce different answers across runs. The model is using probabilities, not a fixed script.

Inference does not change the model

During normal inference, the model’s parameters stay fixed. If a user tells the model a new fact in a prompt, that fact can affect the current response, but inference alone never updates the weights. An application may store conversation history or user memories separately, but that storage is outside the model’s parameters.

This distinction matters:

  • Inference uses the trained model to compute an output.
  • Prompting supplies temporary instructions and context.
  • Retrieval adds external information to the context before inference.
  • Fine-tuning changes model parameters through additional training.

If a support bot answers using a policy document placed in the prompt, the model has not learned that policy forever. The application supplied it for that request.

Why inference matters for engineering

Inference is where a model architecture becomes a production system. The same model can feel fast, slow, reliable, or unpredictable depending on how inference is designed.

Important engineering concerns include:

  • Latency: track time to first token, total response time, and tokens per second.
  • Throughput: batching and continuous batching group active requests so hardware stays busy.
  • Cost: most hosted models charge for input and output tokens.
  • Streaming: responses can be shown token by token as they are generated.
  • Memory: the KV cache, model weights, and batch size all compete for GPU memory.
  • Determinism: lower randomness makes outputs more repeatable, but not always perfectly identical across providers or hardware.
  • Context limits: the prompt, retrieved documents, conversation history, and output all compete for the same context window.
  • Failure handling: applications need timeouts, retries, validation, and fallbacks when output is missing or malformed.

Teams also use inference optimizations. Quantization stores weights with fewer bits to save memory, a tradeoff covered in parameters and weights. Speculative decoding asks a smaller draft model to propose tokens and a larger model to verify them, which can improve speed when the drafts are often right.

For structured tasks, do not rely only on a clever prompt. Use constrained output formats, validators, retries with clear errors, and tests based on real examples.

The key idea

Inference is the runtime process of using a trained model. The model reads tokens, predicts the next token, appends it, and repeats. Prefill, decode, the KV cache, and decoding strategies shape speed and output style, while engineering choices around latency, cost, context, batching, and validation determine whether the model works reliably in a real product.