Definition
RLHF
Reinforcement Learning from Human Feedback is the post-training process that turned raw text predictors into assistants that follow instructions. It runs in three stages: fine-tune on human-written ideal responses, train a reward model from human rankings of candidate outputs, then optimize the language model with PPO to score higher on that reward model. It teaches presentation, not new facts, and it made ChatGPT possible.

Explained in
Chapter 5: Inside the Machine
Everything you need to know about what's happening inside the box. No PhD required.
Related terms