Glossary

Reinforcement Learning from Human Feedback (RLHF)

A training technique that uses human preference judgments to fine-tune AI models toward being more helpful, harmless, and honest — the key method behind how ChatGPT, Claude, and similar assistants behave.


What it means

Reinforcement learning from human feedback (RLHF) is a training technique used to align large language models with human preferences after initial pre-training. It works in three steps:

  1. Supervised fine-tuning (SFT): The model is fine-tuned on a dataset of human-written demonstrations of desired behavior
  2. Reward model training: Human annotators rank multiple model outputs from best to worst; these preferences train a separate “reward model” that scores how good a response is
  3. Reinforcement learning: The main model is optimized using the reward model’s scores as a signal, pushing it toward outputs that humans prefer

RLHF is why instructable models — ChatGPT, Claude, Gemini — follow instructions, decline harmful requests, and maintain conversational tone rather than producing raw language model output. Pre-trained language models without RLHF tend to produce text that continues a prompt statistically rather than responding helpfully.

A related technique, Direct Preference Optimization (DPO), achieves similar results without the separate reward model and has become increasingly common in more recent model development.

Why it matters for researchers

RLHF explains model behavior, not just capability. When a model refuses a request, hedges on a claim, or uses a particular conversational register, this often reflects RLHF training choices rather than the model’s underlying capabilities. Understanding this helps researchers calibrate expectations:

  • Refusals are policy decisions baked in through RLHF, not fundamental limitations of the underlying model architecture
  • Confident tone — RLHF models are trained to sound helpful, which can make them sound more confident than they should be about uncertain claims
  • Consistency issues — because the reward model is itself imperfect, RLHF-trained models can give different quality answers to the same question asked in different ways

Research reproducibility: If you use an LLM in your research workflow, the specific model version matters because RLHF fine-tuning changes between versions. Document the exact model name and version (e.g., claude-sonnet-4-6, gpt-4o-2024-11-20) when reporting results.