Definition

Evaluation (Evals)

Evaluation, or evals, is measuring whether an agent's outputs are good, as distinct from tracing what happened. The conventional approach hand-builds datasets of input, output, and expected output, which goes stale every time the prompt or model changes. The book's alternative uses the system prompt as ground truth, real trajectories as test cases, and specialized scorers across behavior, RAG quality, hallucination, safety, and conversation quality.

Evaluation (Evals) diagram from Intelligence at Scale
Diagram from chapter 13, Production AI: Deployment, Monitoring, and Evaluation

Explained in

Chapter 13: Production AI: Deployment, Monitoring, and Evaluation

The reality of running AI agents. It's nothing like the demo.

Related terms

This is one term. The chapter is the argument.

Intelligence at Scale: 22 chapters, 65,000 words, 80-plus diagrams. Kindle, paperback and hardcover on Amazon.

Buy on Amazon.com
← All terms