Definition
Evaluation (Evals)
Evaluation, or evals, is measuring whether an agent's outputs are good, as distinct from tracing what happened. The conventional approach hand-builds datasets of input, output, and expected output, which goes stale every time the prompt or model changes. The book's alternative uses the system prompt as ground truth, real trajectories as test cases, and specialized scorers across behavior, RAG quality, hallucination, safety, and conversation quality.

Explained in
Chapter 13: Production AI: Deployment, Monitoring, and Evaluation
The reality of running AI agents. It's nothing like the demo.
Related terms