← Back to concepts
9 min read

Attention mechanism

The attention mechanism lets a model compute weighted mixes of information from other token positions. Instead of treating every token as equally important for every other token, it learns mixing weights for each situation.

Those weights are useful, but they are not proof of “importance” or a complete explanation of reasoning. They are learned coefficients inside a larger LLM architecture. Attention helps a model connect related words even when they are far apart in a sentence or document.

A token gathers information from relevant earlier tokens A simplified self-attention example shows the token it as the current query. Its attention weights over selected earlier tokens sum to one: server 0.63, rack 0.21, engineer 0.11, and moved 0.05. The strongest line goes to server, so that value contributes most to the updated representation of it. Illustrative attention weights for the token "it" Weights for this query sum to 1.00; thicker lines mean larger contribution. engineer 0.11 moved 0.05 server 0.63 rack 0.21 it current query Updated representation mostly carries information from "server"
Attention does not choose one word and ignore the rest; it forms a weighted mixture where more relevant tokens contribute more.

A simple example

Consider:

The engineer moved the server into the rack because it was noisy.

To interpret it, a model needs to consider the surrounding words and decide what it refers to. Attention gives the representation for each token a way to use information from other token positions.

Attention does not use a hand-written grammar rule. It learns useful relationships from training examples.

Queries, keys, and values

Most transformer attention is explained with three learned projections. A projection is a learned matrix that turns each input token vector into a new vector for a specific role. The query, key, and value vectors are computed again for each token and each input.

  • A query describes what the current token is looking for.
  • A key describes what each token can be matched on.
  • A value contains the information that can be passed to the current token.

The model compares a query with keys to produce scores. It turns those scores into weights and uses them to combine values. A high weight means that a token contributes more information to the current representation. These learned projections come from the model’s parameters and weights.

The common formula is scaled dot-product attention:

softmax(Q K^T / sqrt(d_k)) V

Read it left to right:

  • Q K^T compares every query with every key to make a score table.
  • / sqrt(d_k) scales the scores so large vectors do not make softmax too peaky.
  • softmax(...) turns each row of scores into weights that sum to 1.
  • Multiplying by V forms a weighted sum of value vectors.

The diagram below uses one row of that score table: the query for it compared with the keys for nearby tokens. Its illustrative weights sum to 1.00, then those weights mix the value vectors into a new embedding-like representation for it.

From query-key scores to a weighted value mixture The query for it is compared with four keys. Example scores of 1.8, 0.7, 0.1, and minus 0.8 are passed through softmax to produce weights 0.63, 0.21, 0.11, and 0.05, which sum to one. The matching values are then combined into one updated token representation. One attention head, simplified Query q("it") Compare with keys KEY SCORE server 1.8 rack 0.7 engineer 0.1 moved -0.8 Softmax weights sum = 1.00 server 0.63 rack 0.21 engineer 0.11 moved 0.05 Weighted sum 0.63 * value(server) + 0.21 * value(rack) + smaller values new vector for "it" Scores and weights are illustrative; the weights for one query sum to 1.
Query-key comparisons decide the weights; the weighted value mixture becomes the next representation for that token.

The names are an analogy, not separate databases. Queries, keys, and values are numerical vectors produced inside the network.

Self-attention and multiple heads

In self-attention, queries, keys, and values all come from the same sequence. Every token can potentially use information from other tokens in that sequence.

That sentence is only true for unmasked, bidirectional attention. Encoder models can let a token use both earlier and later positions. Decoder-only models use a causal mask, so each token can use only itself and earlier positions.

Modern architectures commonly use multi-head attention. Instead of running one large attention calculation, the model splits the hidden vector into several smaller heads, runs attention for all heads in parallel, concatenates their outputs, and projects the result back to the model’s hidden size. Each head can learn a different kind of relationship, such as:

  • Nearby grammatical structure.
  • A reference between a pronoun and a noun.
  • A repeated entity or topic.
  • A pattern in code or a document structure.

The heads’ results are combined before the next part of the layer processes them. Do not overinterpret a single head: the model’s behavior comes from many heads, feed-forward networks, layers, and decoding choices working together.

Masks and positions

An attention mask controls which positions a token may use. In a decoder-only language model, a causal mask prevents a token from looking at future tokens during inference and training. During training, those future tokens are present in the example but hidden by the mask. During generation, they do not exist yet. Without that restriction, training would reveal the answer before the model had to predict it.

In the mask grid below, rows are the token asking the question, called the query. Columns are the tokens being looked at, called keys. A usable cell means “this query may use that key.”

Causal masking allows only current and earlier positions A five-token causal attention mask is shown as a lower-triangular grid. The row for each query token may attend to columns at the same position or earlier positions. Cells to the right are future tokens and are blocked. For example, the token was can attend to The, server, and was, but not to noisy or today. Decoder-only causal mask Rows are query tokens; columns are positions they may use. The server was noisy today The server was noisy today use block use block Rule A row can use cells on or left of its diagonal position. Future columns are blocked.
A causal mask lets each position use the prefix it has seen, but not tokens to its right.

Other masks serve different jobs:

  • Padding masks hide blank padding tokens that are added so examples in a batch have the same length.
  • Cross-attention masks control which decoder positions may look at encoder outputs in encoder-decoder models.
  • Causal masks enforce left-to-right generation.

Attention also needs positional information. Without it, the model would know which tokens are present but not their order. Position signals help distinguish:

The dog chased the cat.
The cat chased the dog.

Models add order in different ways. Some use learned position embeddings. Some use fixed sinusoidal patterns. Many modern decoder models use rotary position embeddings, or RoPE, which rotates query and key vectors in a way that encodes relative position. The practical point is simple: attention compares token vectors, and position signals tell the model where those tokens are in the sequence.

Why attention has a cost

Basic self-attention compares many pairs of token positions. With full attention, the score table grows with the square of sequence length: twice as many tokens means about four times as many query-key scores. This is one reason long-context models need specialized attention variants, careful memory management, or limits on input size from context engineering.

Self-attention pair counts grow with context length A simplified bar chart compares full self-attention pair counts. Four tokens produce sixteen pair checks, eight tokens produce sixty-four pair checks, and sixteen tokens produce two hundred fifty-six pair checks. Each doubling of context length creates four times as many full pair checks. Full self-attention pair checks, simplified Counts shown as n * n for full self-attention before implementation optimizations. 4 tokens 16 pairs 8 tokens 64 pairs 16 tokens 256 pairs longer context Causal masks reduce usable cells, but long contexts still increase attention work.
Attention cost grows with pairs of positions, so longer context can quickly increase memory and compute.

Optimized kernels such as FlashAttention reduce memory traffic and make attention faster on modern GPUs, but they do not remove the quadratic number of query-key scores for full attention. During inference, a KV cache stores earlier keys and values so the model can decode the next token without recomputing the whole prefix.

Architecture variants reduce specific costs:

  • Grouped-query attention and multi-query attention share keys and values across groups of heads, which shrinks the KV cache.
  • Sliding-window attention lets each token attend only to a recent window, which lowers long-context cost but limits direct access to distant tokens.
  • Sparse or block attention patterns compare selected groups instead of every pair.

Attention weights are also not a complete explanation of a model’s reasoning. They show one part of the computation, but a prediction depends on all layers, parameters, and the decoding process.

The key idea

Attention is a learned information-sharing mechanism. It scores relationships between token representations and combines the most relevant information. Queries, keys, values, multiple heads, positional signals, and masks work together to help transformers process context.