Definition
RAG
Retrieval-Augmented Generation gives a model your data at inference time instead of retraining it. Documents are chunked, converted to embeddings, and stored in a vector database; a user's question is embedded the same way, the nearest chunks are retrieved, and they are fed to the model alongside the question. The book's summary: it is a search engine wired into a language model, and it grounds answers in real documents.

Explained in
Chapter 9: What Is an Agent?
The moment you give an LLM access to knowledge or tools, it becomes an agent. And that makes all the difference.
Related terms