Definition
Semantic Caching
Semantic caching returns a stored response when a new question is similar in meaning to one already answered, not just identical. Embed the input, check the cache, serve the hit if fresh, otherwise call the model. Exact-match caching cuts LLM costs 20 to 40 percent for customer-facing agents; semantic caching pushes that to 30 to 50. The hard part is invalidation when the knowledge base changes.
Explained in
Chapter 13: Production AI: Deployment, Monitoring, and Evaluation
The reality of running AI agents. It's nothing like the demo.
Related terms