Reinforcement learning from human feedback, or RLHF, is a post-training method that uses human preferences to shape how a model responds. Instead of defining a perfect answer for every prompt, people compare candidate responses and indicate which one is better. Those comparisons then guide updates to the model’s parameters.
The approach became widely known through its use in assistant-style language models. It is one family of preference-training methods, not a guarantee that every preferred response is correct.
Why use preferences?
Many prompts have several acceptable answers. A useful response may need to be clear, concise, honest about uncertainty, safe, and well formatted. It is difficult to write one exact target response for every situation.
Preference data can express judgments such as:
- Response A follows the user’s instruction better than response B.
- Response A is less likely to invent a fact.
- Response B gives a safer alternative to a harmful request.
- Response A uses a clearer explanation.
These comparisons provide a signal about behavior that ordinary next-token pre-training does not fully specify.
The main stages
Classic RLHF has three stages. First, supervised fine-tuning (SFT) trains a useful instruction-following starting model, as described in the fine-tuning article. Then people provide pairwise comparisons or rankings of candidate answers. Those comparisons train a reward model. Finally, the policy — the model being trained, viewed as choosing output tokens — is optimized against that reward while staying close to the supervised starting model. A KL penalty is the stay-close term that keeps the new policy from drifting too far from the SFT reference model.
1. Supervised fine-tuning
First, the base model is often trained on examples of instructions and desired responses. This supervised fine-tuning, or SFT, gives it a useful starting behavior.
2. Reward model
The model generates several candidate answers for example prompts. Human reviewers usually provide pairwise comparisons, such as chosen versus rejected, or rankings across several answers. A separate reward model learns to score responses in a way that approximates those rankings. Some teams call this a preference model, but reward model is the usual term in classic RLHF.
3. Policy optimization
The language model is treated as the policy that generates actions, where the actions are tokens. A reinforcement-learning algorithm, often PPO in classic RLHF pipelines, updates the policy to produce responses with higher predicted reward. The KL penalty limits how far the policy moves from the starting SFT model, so optimization does not exploit the reward model too aggressively.
In simple terms:
prompt -> candidate response -> reward score -> KL penalty -> model update
RLHF vs DPO and other alternatives
Classic RLHF is powerful, but it is expensive and operationally complex because it trains a reward model and then runs a reinforcement-learning loop. Direct preference optimization (DPO) uses the same preferred/rejected pairs but optimizes the policy directly with a simpler loss. It does not need a separate reward model or PPO loop, which makes it cheaper and easier to run.
Other alternatives change the source of feedback. RLAIF uses AI feedback instead of, or alongside, human feedback. Constitution- or rule-based feedback scores answers against written principles. Verifiable-reward reinforcement learning uses checkable outcomes, such as passing unit tests or producing a correct math answer, which is common for code and math training.
What can go wrong
RLHF optimizes the signals that reviewers and reward models provide. Those signals are incomplete:
- Reviewers may disagree or apply a policy inconsistently, so the reward model learns a blurred signal.
- A reward model can be fooled by fluent but incorrect answers.
- Reward hacking can appear when the model exploits a proxy. For example, if longer answers often receive higher scores, the model may become verbose even when a short answer is better.
- The model may learn sycophancy, agreeing with the user instead of correcting a false premise.
- Safety behavior may become overcautious, causing over-refusal on safe requests.
- Narrow preference data may reduce performance on less common tasks or older capabilities.
This is an example of Goodhart’s law: when a proxy score becomes the target, it can stop measuring the real goal. A high reward does not prove that the response is factually correct or useful in the real world.
Evaluation
Teams should test RLHF models with held-out preference data, factuality checks, safety evaluations, and real user scenarios, ideally through repeatable harness engineering. Human evaluations remain useful because the training reward is only a proxy for user value. Regression tests should check capabilities that must not get worse, such as coding, math, tool use, formatting, and refusal behavior.
Compare against the pre-trained model, the SFT model, and previous preference-tuned checkpoints. If a model scores better on preference data but worse on critical tasks, the reward objective has not captured the full product goal.
The key idea
Classic RLHF shapes a model using human comparisons rather than requiring one perfect answer for every prompt. It can improve helpfulness, safety, and instruction following, but its reward is only a proxy. Careful reviewer guidance and independent evaluation remain essential.