The attention mechanism lets a model compute weighted mixes of information from other token positions. Instead of treating every token as equally important for every other token, it learns mixing weights for each situation.
Those weights are useful, but they are not proof of “importance” or a complete explanation of reasoning. They are learned coefficients inside a larger LLM architecture. Attention helps a model connect related words even when they are far apart in a sentence or document.
A simple example
Consider:
The engineer moved the server into the rack because it was noisy.
To interpret it, a model needs to consider the surrounding words and decide what it refers to. Attention gives the representation for each token a way to use information from other token positions.
Attention does not use a hand-written grammar rule. It learns useful relationships from training examples.
Queries, keys, and values
Most transformer attention is explained with three learned projections. A projection is a learned matrix that turns each input token vector into a new vector for a specific role. The query, key, and value vectors are computed again for each token and each input.
- A query describes what the current token is looking for.
- A key describes what each token can be matched on.
- A value contains the information that can be passed to the current token.
The model compares a query with keys to produce scores. It turns those scores into weights and uses them to combine values. A high weight means that a token contributes more information to the current representation. These learned projections come from the model’s parameters and weights.
The common formula is scaled dot-product attention:
softmax(Q K^T / sqrt(d_k)) V
Read it left to right:
Q K^Tcompares every query with every key to make a score table./ sqrt(d_k)scales the scores so large vectors do not make softmax too peaky.softmax(...)turns each row of scores into weights that sum to 1.- Multiplying by
Vforms a weighted sum of value vectors.
The diagram below uses one row of that score table: the query for it compared with the keys for nearby tokens. Its illustrative weights sum to 1.00, then those weights mix the value vectors into a new embedding-like representation for it.
The names are an analogy, not separate databases. Queries, keys, and values are numerical vectors produced inside the network.
Self-attention and multiple heads
In self-attention, queries, keys, and values all come from the same sequence. Every token can potentially use information from other tokens in that sequence.
That sentence is only true for unmasked, bidirectional attention. Encoder models can let a token use both earlier and later positions. Decoder-only models use a causal mask, so each token can use only itself and earlier positions.
Modern architectures commonly use multi-head attention. Instead of running one large attention calculation, the model splits the hidden vector into several smaller heads, runs attention for all heads in parallel, concatenates their outputs, and projects the result back to the model’s hidden size. Each head can learn a different kind of relationship, such as:
- Nearby grammatical structure.
- A reference between a pronoun and a noun.
- A repeated entity or topic.
- A pattern in code or a document structure.
The heads’ results are combined before the next part of the layer processes them. Do not overinterpret a single head: the model’s behavior comes from many heads, feed-forward networks, layers, and decoding choices working together.
Masks and positions
An attention mask controls which positions a token may use. In a decoder-only language model, a causal mask prevents a token from looking at future tokens during inference and training. During training, those future tokens are present in the example but hidden by the mask. During generation, they do not exist yet. Without that restriction, training would reveal the answer before the model had to predict it.
In the mask grid below, rows are the token asking the question, called the query. Columns are the tokens being looked at, called keys. A usable cell means “this query may use that key.”
Other masks serve different jobs:
- Padding masks hide blank padding tokens that are added so examples in a batch have the same length.
- Cross-attention masks control which decoder positions may look at encoder outputs in encoder-decoder models.
- Causal masks enforce left-to-right generation.
Attention also needs positional information. Without it, the model would know which tokens are present but not their order. Position signals help distinguish:
The dog chased the cat.
The cat chased the dog.
Models add order in different ways. Some use learned position embeddings. Some use fixed sinusoidal patterns. Many modern decoder models use rotary position embeddings, or RoPE, which rotates query and key vectors in a way that encodes relative position. The practical point is simple: attention compares token vectors, and position signals tell the model where those tokens are in the sequence.
Why attention has a cost
Basic self-attention compares many pairs of token positions. With full attention, the score table grows with the square of sequence length: twice as many tokens means about four times as many query-key scores. This is one reason long-context models need specialized attention variants, careful memory management, or limits on input size from context engineering.
Optimized kernels such as FlashAttention reduce memory traffic and make attention faster on modern GPUs, but they do not remove the quadratic number of query-key scores for full attention. During inference, a KV cache stores earlier keys and values so the model can decode the next token without recomputing the whole prefix.
Architecture variants reduce specific costs:
- Grouped-query attention and multi-query attention share keys and values across groups of heads, which shrinks the KV cache.
- Sliding-window attention lets each token attend only to a recent window, which lowers long-context cost but limits direct access to distant tokens.
- Sparse or block attention patterns compare selected groups instead of every pair.
Attention weights are also not a complete explanation of a model’s reasoning. They show one part of the computation, but a prediction depends on all layers, parameters, and the decoding process.
The key idea
Attention is a learned information-sharing mechanism. It scores relationships between token representations and combines the most relevant information. Queries, keys, values, multiple heads, positional signals, and masks work together to help transformers process context.