← Back to concepts
7 min read

Encoder-decoder models

Encoder-decoder models are AI models designed to transform one sequence into another. They are common in tasks like translation, summarization, some speech recognition systems, and structured text generation.

The basic idea is simple: the encoder reads the input, and the decoder writes the output. Modern versions use transformer layers, tokens, embeddings, and attention to connect the two sequences.

A simple example

Imagine translating English to Spanish:

Input:  The meeting starts at five.
Output: La reunion comienza a las cinco.

The Spanish example omits the accent in reunion to keep the code block ASCII-only. The encoder reads the English sentence and builds a sequence of contextual representations, one per input token. The decoder uses that encoded sequence to generate the Spanish sentence.

The model is not copying words one by one. It is learning how the meaning of the input should be expressed in the output format.

The flow through an encoder-decoder model An English input passes through an encoder that reads the full sequence. The encoder produces contextual representations. A decoder repeatedly combines those representations with the Spanish tokens generated so far to produce the next output token. READ ONCE AT INFERENCE Input sequence The meeting starts at five. Encoder reads every input token with surrounding context whole input available Encoded input contextual token representations GENERATE REPEATEDLY Decoder uses encoded input + tokens so far predict next token Output La ... cinco. generated tokens feed the next step
At inference time, the encoder processes the source once; the decoder reuses the encoded token sequence while building the output one token at a time.

What the encoder does

The encoder processes the input sequence. It turns tokens into contextual representations using bidirectional self-attention, so each input token representation can use the surrounding input tokens.

The word “bank” has different meanings in different sentences:

I deposited money at the bank.
I sat near the bank of the river.

An encoder uses the surrounding words to represent the correct meaning. It reads the whole input and produces a sequence of representations that capture the relationships between tokens. It is not compressing the whole source into one single vector.

What the decoder does

The decoder generates the output sequence. It usually writes one token at a time.

At each step, it looks at:

  • What it has already generated.
  • The encoded representation of the input.
  • The task it is trying to complete.

This is useful for tasks where the output length may be different from the input length. A short article may become a one-sentence summary. A sentence in one language may become a longer sentence in another language.

Training and inference look different. During training, encoder-decoder models often use teacher forcing: the decoder receives the correct previous target tokens, shifted right, and learns to predict the next target token at every position. This lets training process target positions in parallel.

During inference, the correct target is not available. The decoder starts from a start-of-sequence token, uses its own previous outputs, and stops when it predicts an end-of-sequence token or reaches a limit. This gap is one reason generation settings and evaluation matter.

Attention between encoder and decoder

In transformer-based encoder-decoder models, the decoder can attend to the encoder’s output. This means that while generating each output token, the decoder can focus on the most relevant parts of the input.

For translation, this helps the model align words and phrases between languages. For summarization, it helps the model focus on the parts of the document that matter for the current sentence.

There are three attention paths:

  1. Encoder self-attention: input tokens attend to other input tokens, usually bidirectionally.
  2. Decoder masked self-attention: output tokens attend only to earlier output tokens.
  3. Cross-attention: decoder queries attend to encoder keys and values, which are the encoded input representations.
Three attention paths in an encoder-decoder transformerThe source input enters the encoder, where bidirectional self-attention mixes input tokens. The target prefix enters the decoder, where masked self-attention can use only earlier output tokens. Cross-attention connects decoder queries to the encoder output sequence, so the decoder can retrieve source information for each generated token.SOURCE SIDEInput tokensThe meeting startsEncoderself-attentionbidirectionalEncoder output sequencekeys and values for sourceTARGET SIDEOutput so farLa reunion ...Decodermasked self-attentionuses earlier outputsNext output tokencincocross-attention
Encoder self-attention reads the source, decoder self-attention tracks the output prefix, and cross-attention connects the two.
Cross-attention while generating one translated token The decoder state after generating La reunion comienza a las forms a query. Cross-attention compares that query with encoded representations for the English source tokens. The strongest illustrative connection is to five, so its encoded information contributes most when the decoder predicts cinco. ENCODER OUTPUT: KEYS AND VALUES The encoded meeting encoded starts encoded at encoded five encoded DECODER STATE: QUERY Output so far La reunion comienza a las ... Cross-attention weight source tokens for this output step strongest illustrative weight Next-token prediction cinco then generation continues The connection weights are illustrative, not measured model behavior.
Cross-attention lets every decoder step retrieve different information from the same encoded input.

Common use cases

Encoder-decoder models are useful when the task is “read this, then produce that.”

Examples include:

  • Translation from one language to another.
  • Summarization of long text.
  • Turning speech into text in systems that use encoder-decoder designs, such as Whisper-style models.
  • Turning natural language into structured queries.
  • Rewriting text in a different style.
  • Generating answers from a provided passage.

These tasks all involve understanding an input and generating a related output.

For translation and summarization, the decoder often uses beam search during inference. Instead of keeping only one partial output, beam search keeps several candidate outputs, extends each one, and chooses the best-scoring completed sequence. This can improve tasks where a globally good output matters more than creative variety. The broader decoding settings are covered in inference.

Encoder-only, decoder-only, and encoder-decoder

Different transformer designs fit different tasks:

Architecture How it works Good for
Encoder-only Reads input and builds representations Classification, extraction, search
Decoder-only Generates text from left to right Chat, completion, code generation
Encoder-decoder Reads input, then generates output Translation, summarization, transformation

Modern chat models are often decoder-only, but encoder-decoder models are still important for many production tasks.

Historically, sequence-to-sequence systems often used recurrent neural networks with attention before transformers became dominant. The modern transformer version keeps the “read, then write” shape but trains much more efficiently on parallel hardware.

Encoder-only, decoder-only, and encoder-decoder information flows Three equal panels compare transformer designs. Encoder-only models read a complete input and return representations or labels. Decoder-only models use a prefix and their previously generated tokens to repeatedly predict the next token. Encoder-decoder models first encode a complete source sequence and then condition a separate decoder on that encoded source while producing a related output. Encoder-only understand a complete input all input tokens visible Encoder stack bidirectional context Label, score, or embedding Decoder-only continue a prefix from left to right prompt + generated tokens so far Decoder stack causal context Next token, then repeat Encoder-decoder transform one sequence into another source input Encoder read output so far Decoder write cross-attention Related output sequence
The architectures share transformer building blocks, but they expose different information to the model and produce different kinds of outputs.

Why engineers should care

Choosing the right architecture affects cost, latency, and quality. If you only need to classify text, an encoder-only model may be enough. If you need open-ended writing, a decoder-only model may fit better. If you need to transform one input into one output, an encoder-decoder model may be a strong choice.

You do not always choose architecture directly when using hosted APIs, but understanding the difference helps you evaluate model behavior and choose the right system around it.

The key idea

Encoder-decoder models split the work into two parts: reading and writing. The encoder understands the input. The decoder generates the output. This design is especially useful when the task transforms one sequence into another.