Encoder-decoder models are AI models designed to transform one sequence into another. They are common in tasks like translation, summarization, some speech recognition systems, and structured text generation.
The basic idea is simple: the encoder reads the input, and the decoder writes the output. Modern versions use transformer layers, tokens, embeddings, and attention to connect the two sequences.
A simple example
Imagine translating English to Spanish:
Input: The meeting starts at five.
Output: La reunion comienza a las cinco.
The Spanish example omits the accent in reunion to keep the code block ASCII-only. The encoder reads the English sentence and builds a sequence of contextual representations, one per input token. The decoder uses that encoded sequence to generate the Spanish sentence.
The model is not copying words one by one. It is learning how the meaning of the input should be expressed in the output format.
What the encoder does
The encoder processes the input sequence. It turns tokens into contextual representations using bidirectional self-attention, so each input token representation can use the surrounding input tokens.
The word “bank” has different meanings in different sentences:
I deposited money at the bank.
I sat near the bank of the river.
An encoder uses the surrounding words to represent the correct meaning. It reads the whole input and produces a sequence of representations that capture the relationships between tokens. It is not compressing the whole source into one single vector.
What the decoder does
The decoder generates the output sequence. It usually writes one token at a time.
At each step, it looks at:
- What it has already generated.
- The encoded representation of the input.
- The task it is trying to complete.
This is useful for tasks where the output length may be different from the input length. A short article may become a one-sentence summary. A sentence in one language may become a longer sentence in another language.
Training and inference look different. During training, encoder-decoder models often use teacher forcing: the decoder receives the correct previous target tokens, shifted right, and learns to predict the next target token at every position. This lets training process target positions in parallel.
During inference, the correct target is not available. The decoder starts from a start-of-sequence token, uses its own previous outputs, and stops when it predicts an end-of-sequence token or reaches a limit. This gap is one reason generation settings and evaluation matter.
Attention between encoder and decoder
In transformer-based encoder-decoder models, the decoder can attend to the encoder’s output. This means that while generating each output token, the decoder can focus on the most relevant parts of the input.
For translation, this helps the model align words and phrases between languages. For summarization, it helps the model focus on the parts of the document that matter for the current sentence.
There are three attention paths:
- Encoder self-attention: input tokens attend to other input tokens, usually bidirectionally.
- Decoder masked self-attention: output tokens attend only to earlier output tokens.
- Cross-attention: decoder queries attend to encoder keys and values, which are the encoded input representations.
Common use cases
Encoder-decoder models are useful when the task is “read this, then produce that.”
Examples include:
- Translation from one language to another.
- Summarization of long text.
- Turning speech into text in systems that use encoder-decoder designs, such as Whisper-style models.
- Turning natural language into structured queries.
- Rewriting text in a different style.
- Generating answers from a provided passage.
These tasks all involve understanding an input and generating a related output.
For translation and summarization, the decoder often uses beam search during inference. Instead of keeping only one partial output, beam search keeps several candidate outputs, extends each one, and chooses the best-scoring completed sequence. This can improve tasks where a globally good output matters more than creative variety. The broader decoding settings are covered in inference.
Encoder-only, decoder-only, and encoder-decoder
Different transformer designs fit different tasks:
| Architecture | How it works | Good for |
|---|---|---|
| Encoder-only | Reads input and builds representations | Classification, extraction, search |
| Decoder-only | Generates text from left to right | Chat, completion, code generation |
| Encoder-decoder | Reads input, then generates output | Translation, summarization, transformation |
Modern chat models are often decoder-only, but encoder-decoder models are still important for many production tasks.
Historically, sequence-to-sequence systems often used recurrent neural networks with attention before transformers became dominant. The modern transformer version keeps the “read, then write” shape but trains much more efficiently on parallel hardware.
Why engineers should care
Choosing the right architecture affects cost, latency, and quality. If you only need to classify text, an encoder-only model may be enough. If you need open-ended writing, a decoder-only model may fit better. If you need to transform one input into one output, an encoder-decoder model may be a strong choice.
You do not always choose architecture directly when using hosted APIs, but understanding the difference helps you evaluate model behavior and choose the right system around it.
The key idea
Encoder-decoder models split the work into two parts: reading and writing. The encoder understands the input. The decoder generates the output. This design is especially useful when the task transforms one sequence into another.