← intelligenceatscale.ai

Glossary

The vocabulary of building with LLMs.

68 terms from the book, defined the way the chapters define them, each linked to the chapter and the diagram that explains it.

A

A2A (Agent-to-Agent Protocol)

Ch. 21

Agent-to-Agent Protocol is Google's standard for agents in different organizations to discover and work with each other, the HTTP of the agent economy.
Agent

Ch. 9

An agent is an LLM given access to external knowledge through RAG, the ability to take actions through tool calling, or both.
Agent Loop

Ch. 10

The agent loop is the core execution pattern under every framework: send the model a message plus tool definitions, let it decide whether to call a tool, execute the tool, append the result, and call the model again until it produces a final answer.
Agent Memory

Ch. 12

Agent memory is everything competing for space in the context window.
Agent Spawning

Ch. 21

Agent spawning is a supervisor creating specialists on the fly instead of routing to pre-built ones.
Agentic RAG

Ch. 9

Agentic RAG puts the agent in charge of retrieval instead of running one similarity search and hoping.
AGI

Ch. 22

Artificial General Intelligence has no agreed definition, which the book argues matters more than any timeline.
Autonomy Spectrum

Ch. 21

The autonomy spectrum describes how much an agent does without a person.

B

Backpropagation

Ch. 5

Backpropagation is the 1986 algorithm that still trains every neural network today.
Bounded Autonomy

Ch. 13

Bounded autonomy assigns each action a risk tier that decides how much the agent may do alone.

C

Chain Accuracy

Ch. 11

Chain accuracy is the multiplication that breaks multi-agent systems.
Chain-of-Thought

Ch. 8

Chain-of-thought prompting asks the model to reason step by step before answering, in its simplest form by appending "Let's think step by step." Because tokens are generated sequentially, the intermediate reasoning becomes context for the final answer, effectively a scratchpad.
Circuit Breaker

Ch. 13

A circuit breaker stops an agent stuck in a loop from burning tokens: a tool errors, the agent retries, same error, twenty iterations later the user is staring at a spinner.
Context Engineering

Ch. 8

Context engineering is deciding what goes into the context window, how it is structured, and where each piece sits.
Context Stack

Ch. 15

The context stack is the onboarding material that makes a coding agent expert at your codebase.
Context Window

Ch. 8

The context window is everything the model can see for a single request: system prompt, conversation history, retrieved documents, tool schemas, and tool results.
CUDA

Ch. 6

CUDA is NVIDIA's 2006 programming platform that let ordinary C-style code run on GPUs instead of disguising math as graphics operations.

D

Decoder-Only Transformer

Ch. 3

A decoder-only Transformer uses just the generating half of the original architecture, reading text left to right and predicting the next token.
Diffusion Model

Ch. 18

A diffusion model learns to reverse noise.
Diffusion Transformer (DiT)

Ch. 18

A Diffusion Transformer, or DiT, replaces the U-Net backbone of early diffusion models with a Transformer.

E

Embedding

Ch. 5

An embedding is a vector of hundreds or thousands of numbers that represents a token's meaning as a position in high-dimensional space.
Emergent Abilities

Ch. 3

Emergent abilities are capabilities that appear once a model reaches a certain scale without anyone training for them explicitly.
Evaluation (Evals)

Ch. 13

Evaluation, or evals, is measuring whether an agent's outputs are good, as distinct from tracing what happened.

F

Few-Shot Prompting

Ch. 8

Few-shot prompting means showing the model three to five examples of input and desired output so it mimics the pattern rather than following abstract rules.
Fine-Tuning

Ch. 7

Fine-tuning continues training a pre-trained model on your own examples so its weights adapt to a specific task, style, or vocabulary.

G

GAN

Ch. 18

A Generative Adversarial Network trains two networks against each other: a Generator turns random noise into images while a Discriminator judges real from fake, each improving until the fakes pass, like a counterfeiter versus a detective.
Groundedness

Ch. 9

Groundedness is whether a response is actually supported by the documents retrieved for it, rather than blended with the model's training data.

H

Hallucination

Ch. 5

A hallucination is a confident, fluent statement that is false: invented citations, made-up statistics, misattributed quotes.
Handoff

Ch. 11

A handoff is the swarm-style mechanism, popularized by OpenAI's Swarm framework, where an agent does not call another agent but becomes one: it returns a handoff and the framework transfers control, context, and conversation history to the target.
Huang's Law

Ch. 6

Huang's Law is Jensen Huang's claim that AI computing performance per dollar improves far faster than Moore's Law.
Human-in-the-Loop

Ch. 10

Human-in-the-loop means pausing an agent before a consequential action so a person can approve, reject, or edit it.

K

Kill Switch

Ch. 11

Kill switches are hard limits that stop a multi-agent system from spiraling.

L

LangGraph

Ch. 10

LangGraph models an agent as a directed graph with typed state rather than a linear loop.
LLM Wrapper Trap

Ch. 14

The LLM wrapper trap is building a thin product, model plus UI plus system prompt, that dies the moment the model provider ships your feature as a default.
LLM-as-a-Judge

Ch. 13

LLM-as-a-judge uses one model to score another model's output.
Long-Term Memory

Ch. 12

Long-term memory is what an agent retains across conversations, and the book splits it three ways.

M

Matrix Multiplication

Ch. 5

A neuron computes a weighted sum of its inputs, which is a dot product.
MCP (Model Context Protocol)

Ch. 12

Model Context Protocol is Anthropic's open standard for connecting AI models to tools and data sources, the USB-C for AI.
Model Routing

Ch. 13

Model routing sends each query to the cheapest model that can handle it well.
Multi-Agent System

Ch. 11

A multi-agent system splits work across several specialized agents instead of one overloaded prompt.

O

Observability

Ch. 13

Observability is the ability to explain why an agent did what it did, not just that a metric moved.

P

Positional Encoding

Ch. 3

Because self-attention processes all tokens simultaneously, a Transformer has no built-in sense of word order; "dog bites man" and "man bites dog" would look identical.
Prompt Injection

Ch. 12

Prompt injection is an attack where malicious instructions are hidden in data the agent processes, such as an email or web page, and the model follows them as if they came from its operator.

R

RAG

Ch. 9

Retrieval-Augmented Generation gives a model your data at inference time instead of retraining it.
ReAct

Ch. 8

ReAct, short for Reasoning plus Acting, is the loop most production agents run: the model thinks about what it needs, takes an action such as a tool call, observes the result, and repeats until it can answer.
Remediation Loop

Ch. 13

Remediation, or the detection-to-correction loop, closes the gap between knowing an agent is failing and fixing it.
RLHF

Ch. 5

Reinforcement Learning from Human Feedback is the post-training process that turned raw text predictors into assistants that follow instructions.

S

Scaling Laws

Ch. 7

Scaling laws are the empirical finding that model performance improves predictably, as a power law, with more parameters, more data, and more compute.
Self-Attention

Ch. 3

Self-attention lets every token in a sequence look at every other token at once and decide which ones matter.
Semantic Caching

Ch. 13

Semantic caching returns a stored response when a new question is similar in meaning to one already answered, not just identical.
Shadow Mode

Ch. 13

Shadow mode runs the agent alongside humans before launch: both handle the same requests, but only the human's response reaches the customer.
Span

Ch. 13

A span is one individual action inside a trace: an LLM call, a tool call, a RAG retrieval, a guardrail check.
State Space Model (Mamba)

Ch. 22

A state space model such as Mamba processes a sequence as a signal flowing through a dynamical system rather than comparing every token to every other.
Supervisor Pattern

Ch. 11

The supervisor pattern uses a central router agent that reads each request, delegates it to a specialist, reviews the output, and decides what happens next.
System Prompt

Ch. 8

The system prompt is the block of text that tells a model who it is, what it does, how it behaves, and what it must never do.
System Prompt as Ground Truth

Ch. 13

This is the book's central evaluation idea: the system prompt already specifies what the agent should do, so use it as the spec instead of hand-labeling test cases.

T

Temperature

Ch. 5

Temperature is the randomness dial applied when picking the next token from the model's probability distribution.
Temporal Coherence

Ch. 20

Temporal coherence is making consecutive video frames belong together: no flicker, no morphing objects, no faces that subtly reshape.
Test-Time Compute

Ch. 7

Test-time compute is the third scaling frontier: letting a model think longer at inference time instead of only making it bigger or training it more.
Tokenization

Ch. 5

Tokenization is the step that turns text into numbers a model can process.
Tool Calling

Ch. 9

Tool calling, also called function calling, is how an agent acts on the world.
Tool Description

Ch. 12

A tool description is the text in a function schema that tells the model when a tool is relevant and how to format its arguments.
Trace

Ch. 13

A trace is one complete agent run: the user's input, every reasoning step, retrieval, tool call, and the final answer, captured as a single top-level record.
Trajectory

Ch. 13

A trajectory is the full execution path an agent actually took for one request: every reasoning step, tool call, retrieval, and response, in order.
Transformer

Ch. 3

The Transformer is the neural network architecture introduced by Google in 2017 that replaced recurrence with attention.

V

Vector Database

Ch. 9

A vector database stores embeddings and answers nearest-neighbor queries: given a query vector, return the stored chunks closest to it in meaning.
Vibe Coding

Ch. 15

Vibe coding is Karpathy's term for trusting AI enough to move fast and iterate, which the internet distorted into shipping unreviewed output.
Voice AI Pipeline

Ch. 19

A voice agent chains three models: speech-to-text transcribes the caller, an LLM reasons and drafts a reply, and text-to-speech renders it as audio.
Buy on Amazon.com

Every term here has a chapter behind it. 22 chapters, 80-plus diagrams.