Definition
LLM-as-a-Judge
LLM-as-a-judge uses one model to score another model's output. It works because verification is easier than generation: proofreading is easier than writing, the P versus NP intuition. MT-Bench showed GPT-4 agreeing with human evaluators over 80 percent of the time, matching human-to-human agreement. The key constraint is that the judge only answers narrow questions against instructions or documents, which collapses its own hallucination surface.
Explained in
Chapter 13: Production AI: Deployment, Monitoring, and Evaluation
The reality of running AI agents. It's nothing like the demo.
Related terms