Definition

LLM-as-a-Judge

LLM-as-a-judge uses one model to score another model's output. It works because verification is easier than generation: proofreading is easier than writing, the P versus NP intuition. MT-Bench showed GPT-4 agreeing with human evaluators over 80 percent of the time, matching human-to-human agreement. The key constraint is that the judge only answers narrow questions against instructions or documents, which collapses its own hallucination surface.

Explained in

Chapter 13: Production AI: Deployment, Monitoring, and Evaluation

The reality of running AI agents. It's nothing like the demo.

Related terms

This is one term. The chapter is the argument.

Intelligence at Scale: 22 chapters, 65,000 words, 80-plus diagrams. Kindle, paperback and hardcover on Amazon.

Buy on Amazon.com
← All terms