← Back to concepts
7 min read

Reinforcement learning from human feedback (RLHF)

Reinforcement learning from human feedback, or RLHF, is a post-training method that uses human preferences to shape how a model responds. Instead of defining a perfect answer for every prompt, people compare candidate responses and indicate which one is better. Those comparisons then guide updates to the model’s parameters.

The approach became widely known through its use in assistant-style language models. It is one family of preference-training methods, not a guarantee that every preferred response is correct.

Why use preferences?

Many prompts have several acceptable answers. A useful response may need to be clear, concise, honest about uncertainty, safe, and well formatted. It is difficult to write one exact target response for every situation.

Preference data can express judgments such as:

  • Response A follows the user’s instruction better than response B.
  • Response A is less likely to invent a fact.
  • Response B gives a safer alternative to a harmful request.
  • Response A uses a clearer explanation.

These comparisons provide a signal about behavior that ordinary next-token pre-training does not fully specify.

The main stages

Classic RLHF has three stages. First, supervised fine-tuning (SFT) trains a useful instruction-following starting model, as described in the fine-tuning article. Then people provide pairwise comparisons or rankings of candidate answers. Those comparisons train a reward model. Finally, the policy — the model being trained, viewed as choosing output tokens — is optimized against that reward while staying close to the supervised starting model. A KL penalty is the stay-close term that keeps the new policy from drifting too far from the SFT reference model.

The classic RLHF pipeline from base model to optimized policy The diagram shows a base model first going through supervised fine-tuning, abbreviated SFT, to become an SFT reference model. The reference model generates candidate answers for prompts. Human reviewers compare candidates. Those comparisons train a reward model. The reward model sends a direct reward signal to policy optimization, while a separate KL penalty keeps the policy close to the SFT reference model. The result is an RLHF-tuned policy. STAGE 1 - instruction tuning Base model from pre-training SFT model reference policy STAGE 2 - reward model Prompts example tasks Candidates A, B, C answers Human comparisons which answer is better? Reward model scores responses STAGE 3 - policy optimization Optimize policy e.g. PPO update reward signal KL penalty keeps it close RLHF-tuned policy model The reward model pulls the policy toward preferred responses; the KL penalty discourages drifting too far from the SFT reference model.
RLHF has two learned pieces: a reward model trained from human comparisons and a policy model optimized against that reward with a stay-close penalty.

1. Supervised fine-tuning

First, the base model is often trained on examples of instructions and desired responses. This supervised fine-tuning, or SFT, gives it a useful starting behavior.

2. Reward model

The model generates several candidate answers for example prompts. Human reviewers usually provide pairwise comparisons, such as chosen versus rejected, or rankings across several answers. A separate reward model learns to score responses in a way that approximates those rankings. Some teams call this a preference model, but reward model is the usual term in classic RLHF.

Training a reward model from preference comparisons For the same prompt, answer A is grounded and concise while answer B is fluent but invents a number. A reviewer marks A as preferred. Many such pairs become comparison data. The reward model learns to assign a higher illustrative score to A, 0.82, than to B, 0.24, and can then score new candidates even when there is no human in the loop for that prompt. Prompt "Summarize the outage for the support team." Answer A Grounded, concise summary reviewer preferred Answer B Fluent, invents a number reviewer rejected Comparison data A better than B repeated across many prompts Reward model scores candidates later A 0.82 B 0.24 Illustrative scores, not truth. The reward model learns reviewer preferences; it does not independently prove that an answer is factually correct.
A reward model converts many human comparisons into a learned scoring function that can guide later policy updates.

3. Policy optimization

The language model is treated as the policy that generates actions, where the actions are tokens. A reinforcement-learning algorithm, often PPO in classic RLHF pipelines, updates the policy to produce responses with higher predicted reward. The KL penalty limits how far the policy moves from the starting SFT model, so optimization does not exploit the reward model too aggressively.

In simple terms:

prompt -> candidate response -> reward score -> KL penalty -> model update
Policy optimization with reward and KL penalty A prompt goes to the current policy, which samples a candidate response. The candidate response goes to the reward model. The same prompt also goes around the main path to the frozen SFT reference model. A KL penalty compares the current policy with the reference model. The objective combines reward minus KL penalty, then a PPO-style update loops around the outside to adjust the current policy. No arrows pass through boxes; the reference model stays frozen. Prompt training task Current policy samples response updates each step Candidate response tokens Reward model predicted preference reward = +0.74 SFT reference frozen model distance to policy KL penalty stay close Objective reward minus KL penalty PPO-style update adjusts the current policy Reward pushes behavior toward reviewer preferences; KL keeps optimization from exploiting the reward model too aggressively.
Policy optimization balances two clean signals: higher reward from the reward model and a KL penalty that keeps the policy close to the SFT reference.

RLHF vs DPO and other alternatives

Classic RLHF is powerful, but it is expensive and operationally complex because it trains a reward model and then runs a reinforcement-learning loop. Direct preference optimization (DPO) uses the same preferred/rejected pairs but optimizes the policy directly with a simpler loss. It does not need a separate reward model or PPO loop, which makes it cheaper and easier to run.

Classic RLHF compared with DPO Two side-by-side paths start from preference pairs. Classic RLHF trains a reward model, then uses PPO with a KL penalty to update the policy. DPO sends the same preferred and rejected pairs directly into a DPO loss that updates the policy without a separate reward model or PPO loop. Classic RLHF Preference pairs chosen vs rejected Reward model learns scores PPO + KL RL update loop Updated policy reward-optimized DPO Preference pairs same data type DPO loss no reward model Updated policy directly tuned
DPO removes the reward-model-and-PPO loop, but it still depends on high-quality preference pairs.

Other alternatives change the source of feedback. RLAIF uses AI feedback instead of, or alongside, human feedback. Constitution- or rule-based feedback scores answers against written principles. Verifiable-reward reinforcement learning uses checkable outcomes, such as passing unit tests or producing a correct math answer, which is common for code and math training.

What can go wrong

RLHF optimizes the signals that reviewers and reward models provide. Those signals are incomplete:

  • Reviewers may disagree or apply a policy inconsistently, so the reward model learns a blurred signal.
  • A reward model can be fooled by fluent but incorrect answers.
  • Reward hacking can appear when the model exploits a proxy. For example, if longer answers often receive higher scores, the model may become verbose even when a short answer is better.
  • The model may learn sycophancy, agreeing with the user instead of correcting a false premise.
  • Safety behavior may become overcautious, causing over-refusal on safe requests.
  • Narrow preference data may reduce performance on less common tasks or older capabilities.

This is an example of Goodhart’s law: when a proxy score becomes the target, it can stop measuring the real goal. A high reward does not prove that the response is factually correct or useful in the real world.

Evaluation

Teams should test RLHF models with held-out preference data, factuality checks, safety evaluations, and real user scenarios, ideally through repeatable harness engineering. Human evaluations remain useful because the training reward is only a proxy for user value. Regression tests should check capabilities that must not get worse, such as coding, math, tool use, formatting, and refusal behavior.

Compare against the pre-trained model, the SFT model, and previous preference-tuned checkpoints. If a model scores better on preference data but worse on critical tasks, the reward objective has not captured the full product goal.

The key idea

Classic RLHF shapes a model using human comparisons rather than requiring one perfect answer for every prompt. It can improve helpfulness, safety, and instruction following, but its reward is only a proxy. Careful reviewer guidance and independent evaluation remain essential.