A recurrent neural network, or RNN, is a neural network designed to process a sequence one step at a time. At each step, it reads the current input and combines it with a hidden state that carries information from earlier steps.
This made RNNs useful for language, speech, time series, and other ordered data before transformers became dominant.
How recurrence works
Imagine processing a sentence from left to right. The RNN updates its hidden state for each token:
token 1 + state 0 -> state 1
token 2 + state 1 -> state 2
token 3 + state 2 -> state 3
The current state is a compact summary of what the network has seen so far. It is a fixed-size learned vector, not stored text. The starting state h0 is often all zeros or a learned initial vector. The model can use the current state to classify the sequence or predict the next item.
The same weights are reused at every step. This lets one network handle sequences of different lengths without creating a separate set of weights for every position.
A simple RNN update is often written like this:
h_t = tanh(W_x x_t + W_h h_(t-1) + b)
In plain words: the new hidden state h_t comes from the current input vector x_t, the previous hidden state h_(t-1), learned weights W_x and W_h, and a bias b. The tanh function squashes the result into a stable range. The same weights are reused at every time step.
The memory problem
The hidden state has limited capacity. As a sequence becomes longer, important information from the beginning can be diluted by later updates. Because the state vector has a fixed size, it must compress the past into a lossy summary.
During training, RNNs use backpropagation through time. The loop is unrolled into a long chain, then trained like a deep network. For long sequences, teams often use truncated BPTT, which backpropagates through only a fixed number of recent steps to save memory and reduce instability.
The model’s error signal must travel backward through many time steps. Repeated multiplication can make that signal become extremely small, a problem called the vanishing gradient. If the signal becomes too small, early steps receive almost no useful update. The 0.6 multiplier below is illustrative; real behavior depends on the learned weights and activation functions.
Very large repeated multiplications can create exploding gradients, where updates become unstable. A common defense is gradient clipping, which caps the update size before it damages training.
Variants such as LSTM and GRU add gates that help decide what to keep, update, or forget. They can remember longer relationships than a simple RNN, but they still process a sequence step by step.
An LSTM, or long short-term memory network, keeps a separate cell state alongside the hidden state. Gates decide what to forget, what new information to write, and what part of the cell state to expose as the next hidden state. A GRU, or gated recurrent unit, is a simpler gated design that combines some of those decisions into two main gates.
RNNs versus transformers
RNNs can also be bidirectional: one RNN reads left to right, another reads right to left, and their states are combined. This helps understanding tasks where the whole input is available, but it cannot be used the same way for left-to-right generation.
RNNs and transformers both process sequences, but they organize computation differently:
| Model | Main approach | Practical effect |
|---|---|---|
| RNN | Carry a hidden state from one step to the next | Natural sequence order, but limited parallelism |
| Transformer | Use attention between token positions | Better parallel training and direct long-range connections |
Because an RNN must usually finish one step before starting the next, it is harder to train efficiently on large datasets. A transformer can process many positions in parallel during training using attention. Transformer generation is still one token at a time during inference, but training is much more parallel.
RNNs can still be a good fit for small streaming systems, compact time-series models, or devices with strict resource limits. The best architecture depends on the task and constraints.
Historically, many sequence-to-sequence translation systems used RNN encoders and decoders. Attention was first added to help those decoders look back at the most relevant source states, a pattern that later became central to encoder-decoder models. Some modern recurrent-style and state-space ideas are also returning for long sequences, usually to reduce the memory cost of full attention.
The key idea
An RNN processes ordered data by carrying a hidden state through a sequence. This idea provides useful memory, but long-range learning and sequential computation are difficult. LSTMs and GRUs improve the design with gates, while transformers use attention to connect positions more directly.