← Back to concepts
10 min read

What is a large language model?

A large language model, or LLM, is an AI model trained to work with language. You give it text, and it predicts useful text in response. Chat assistants, coding assistants, summarizers, and many document tools are built on this basic capability.

The core idea is simple: an LLM predicts what should come next. Repeated many times, next-token prediction can produce paragraphs, code, answers, translations, plans, and conversations.

What makes it large?

An LLM is large in two main ways:

  • Large training data: it learns from huge collections of text, code, and other written material, often measured in trillions of tokens.
  • Large parameter count: it contains many learned numerical values, called parameters or weights, that shape its predictions. Modern LLMs typically have billions of them.

Those learned values are not a database of copied pages. They are patterns spread across the model’s parameters. The model uses those patterns to estimate likely continuations for the input it receives.

Next-token prediction, step by step

A token is a small chunk of text, such as a word, part of a word, or a punctuation mark. At each step, the model looks at every token in the input so far and assigns a probability to every token it knows. Then one token is picked, added to the text, and the whole process runs again.

Here is an illustrative example:

Input so far:  "The capital of France is"

Model scores for the next token (illustrative):
  " Paris"     0.82
  " a"         0.06
  " the"       0.04
  " located"   0.03
  (all others) 0.05

Picked: " Paris"
New input:     "The capital of France is Paris"  -> predict again

This repeating cycle is called autoregressive generation: each new token becomes part of the input for the next prediction. It continues until the model produces a special end token or the application hits a length limit. It is also why answers can stream onto the screen word by word instead of appearing all at once.

The next-token generation loop The context so far, The capital of France is, goes into the LLM. The LLM scores every possible next token. In this illustrative example, Paris gets probability 0.82, a gets 0.06, the gets 0.04, located gets 0.03, and all other tokens share 0.05. One token is picked, Paris, and appended to the context. The loop then repeats until an end token or a length limit is reached. CONTEXT SO FAR "The capital of France is" LLM scores every possible token Next-token probabilities " Paris" 0.82 " a" 0.06 " the" 0.04 " located" 0.03 all others 0.05 Pick one token " Paris" sampling settings control the choice append " Paris" to the context, then predict again Stops at an end token or a length limit. Probabilities are illustrative.
A long answer is many small predictions: score every token, pick one, append it, and repeat.

The model does not always pick the highest-probability token. Settings such as temperature control how adventurous the choice is. Low temperature makes output more predictable; higher temperature makes it more varied. The inference article covers these settings in detail.

How an LLM reads text

An LLM does not read words exactly the way people do. First, a tokenizer splits text into tokens and maps each one to a numeric token ID. Inside the model, each ID becomes a list of numbers called an embedding. Those numbers then pass through a stack of transformer layers that let every token gather information from the tokens before it. The last layer produces a score for every possible next token.

From text to next-token scores A left-to-right pipeline. The text Summarize this ticket goes into the tokenizer, which splits it into tokens. The tokens become illustrative numeric token IDs. Inside the model, each ID becomes an embedding vector, and the vectors pass through many repeated transformer layers that output scores for the next token. The tokenizer prepares the input; the embeddings and transformer layers are the model itself. PREPARING THE INPUT INSIDE THE MODEL Text "Summarize this ticket" Tokenizer splits text into tokens Token IDs [9534, 553, 4876, 11244] Embeddings each ID becomes a vector Transformer layers x N mix context, score next token Token IDs are illustrative; real IDs depend on the tokenizer.
The tokenizer turns text into numbers; the model turns those numbers into a score for every possible next token.

You do not need to understand every layer to build with LLMs. The practical point is that everything the model sees, including instructions, documents, and its own previous output, is processed as tokens.

Training and post-training

Most LLMs go through more than one learning stage.

Pre-training teaches broad language patterns. The model sees a huge amount of text and learns to predict the next token. This stage gives it general ability with grammar, facts, code patterns, and reasoning-like behavior. The result is a base model: good at continuing text, but not yet good at following instructions.

Post-training makes the model more useful as an assistant. It may learn from instruction examples, human preferences, safety rules, and task-specific demonstrations. This stage turns “continue this text” into “answer this request helpfully and safely.”

After training, the model’s parameters are usually fixed during normal use. A prompt can guide the model for one request, but it does not permanently teach the model new information.

Training changes the weights; using the model does not The top row shows training, which changes the model weights. Pre-training on huge amounts of text produces a base model. Post-training on instructions, preferences, and safety examples produces an assistant model. The assistant is deployed with frozen weights. The bottom row shows normal use, which does not change the weights. A prompt plus context goes into the deployed model, and a response comes out for that one request. TRAINING (CHANGES THE WEIGHTS) Pre-training predict the next token on huge text collections result: base model Post-training instructions, preferences, safety examples result: assistant model Deployed model weights are frozen during normal use USE (DOES NOT CHANGE THE WEIGHTS) A prompt guides one request. It does not retrain the model. Prompt + context question, docs, history Response for this request only
Training shapes the weights once; every later request uses those same frozen weights plus whatever context the prompt supplies.

To change what a model knows or how it behaves over the long term, teams either supply better context at request time or train the model further with fine-tuning.

What prompts and context do

A prompt is the input you send to the model. It may include the user’s question, system instructions, examples, retrieved documents, conversation history, or tool results. Writing clear prompts is the focus of prompt engineering.

The model can only use what fits in its context window, the maximum number of tokens it can process at once, plus what it learned during training. If important information is missing from the prompt, the model may guess. Choosing what goes into that window is called context engineering.

For example, a support tool might send this:

System: You summarize support tickets in 3 bullet points.
        Only use facts from the ticket.

Ticket #4812:
Customer says the export button does nothing on Safari 17.
Works in Chrome. Started after Tuesday's release. Customer is on the Pro plan.

Task: Summarize this ticket.

The model knows nothing about ticket #4812 from training. Everything it needs comes from the prompt, which is exactly why good context matters as much as the model itself.

This is also why AI applications often add surrounding systems:

  • Search or retrieval to provide relevant facts.
  • Tools for calculations, database reads, or external actions.
  • Validation to check whether output follows the expected format.
  • Monitoring to track quality, cost, and failures.
An LLM as one part of an application A user sends a request to the application and receives an answer back. Inside the application, retrieval supplies relevant facts and tools supply calculations or external data to the LLM. The LLM generates a draft. Validation checks format and safety before the answer is returned, and monitoring records quality, cost, and errors. The LLM sits in the middle, but the surrounding application parts make the result reliable. User asks, gets an answer Your application Retrieval relevant facts Tools math, search, APIs LLM generates a draft Validation format + safety checks Monitoring quality, cost, errors checked answer returns to the user
The LLM generates the draft, but retrieval, tools, validation, and monitoring are what make the product dependable.

The LLM is important, but it is only one part of a reliable product.

Kinds of LLMs you will meet

The same core idea shows up in several forms:

  • Base models are straight out of pre-training. They continue text well but do not reliably follow instructions.
  • Instruction-tuned or chat models have been post-trained to act as assistants. Most products use these.
  • Reasoning models are post-trained to write out intermediate thinking tokens before the final answer. They are often better at multi-step problems, but they use more tokens and take longer.
  • Multimodal models accept images, audio, or other inputs alongside text.
  • Closed and open-weight models: some models are only available through a provider’s API, while others publish their weights so you can run them yourself.

What LLMs do well

LLMs are useful for tasks that involve language patterns:

  • Drafting and rewriting text.
  • Summarizing documents or conversations.
  • Explaining concepts in different levels of detail.
  • Generating and reviewing code.
  • Extracting structure from messy text.
  • Brainstorming options or first drafts.

They are especially helpful when the task has ambiguity and the user benefits from a flexible answer.

Common limitations

LLMs also have important limits:

  • Hallucination: they can produce false information that sounds confident, because they generate likely-sounding text rather than looking facts up.
  • Stale knowledge: they only know what was in their training data, up to a knowledge cutoff date. Recent events and private facts must be supplied in context.
  • Weak exactness: arithmetic, counting letters, and strict logic can fail without tools or checks. Tokenization is part of the reason: the model sees chunks, not individual characters.
  • Context limits: long documents may not fit, or the model may miss details buried in the middle of them.
  • Non-determinism: the same prompt can produce different outputs depending on sampling settings.
  • Prompt injection: text inside a document or web page can try to override your instructions, because the model sees instructions and data as the same kind of input.

These limits do not make LLMs useless. They mean engineers should design systems that verify important claims, retrieve source material, constrain risky actions, and escalate to a person when confidence is low.

Common misconceptions

  • “It looks things up.” A plain LLM does not search the web or a database. It generates from learned patterns plus the prompt. Search only happens if the application adds a search tool or retrieval step.
  • “It learns from my conversation.” The weights do not change when you chat. A product may store notes or history and feed them back as context, but that is the application, not the model, remembering.
  • “The same question always gets the same answer.” Output depends on sampling settings, context, and the exact wording of the prompt.
  • “A confident answer is a correct answer.” Fluency and accuracy are different things. Check important claims against sources.

The key idea

An LLM is a trained prediction system for language. It produces answers one token at a time, using patterns learned during training and whatever information is in the prompt. In real applications, the model works best when paired with good prompts, relevant context, tools, validation, and clear product boundaries.