Definition

Semantic Caching

Semantic caching returns a stored response when a new question is similar in meaning to one already answered, not just identical. Embed the input, check the cache, serve the hit if fresh, otherwise call the model. Exact-match caching cuts LLM costs 20 to 40 percent for customer-facing agents; semantic caching pushes that to 30 to 50. The hard part is invalidation when the knowledge base changes.

Explained in

Chapter 13: Production AI: Deployment, Monitoring, and Evaluation

The reality of running AI agents. It's nothing like the demo.

Related terms

This is one term. The chapter is the argument.

Intelligence at Scale: 22 chapters, 65,000 words, 80-plus diagrams. Kindle, paperback and hardcover on Amazon.

Buy on Amazon.com
← All terms