Definition
System Prompt as Ground Truth
This is the book's central evaluation idea: the system prompt already specifies what the agent should do, so use it as the spec instead of hand-labeling test cases. Compare the agent's actual trajectory against the prompt using many specialized scorers. It is control theory applied to LLMs: the prompt is the set point, continuous evaluation is the measurement, and correction adjusts drift. It also scales to agents that spawn agents.

Explained in
Chapter 13: Production AI: Deployment, Monitoring, and Evaluation
The reality of running AI agents. It's nothing like the demo.
Related terms