Definition

RLHF

Reinforcement Learning from Human Feedback is the post-training process that turned raw text predictors into assistants that follow instructions. It runs in three stages: fine-tune on human-written ideal responses, train a reward model from human rankings of candidate outputs, then optimize the language model with PPO to score higher on that reward model. It teaches presentation, not new facts, and it made ChatGPT possible.

RLHF diagram from Intelligence at Scale
Diagram from chapter 5, Inside the Machine

Explained in

Chapter 5: Inside the Machine

Everything you need to know about what's happening inside the box. No PhD required.

Related terms

This is one term. The chapter is the argument.

Intelligence at Scale: 22 chapters, 65,000 words, 80-plus diagrams. Kindle, paperback and hardcover on Amazon.

Buy on Amazon.com
← All terms