Agent-to-Agent Protocol is Google's standard for agents in different organizations to discover and work with each other, the HTTP of the agent economy.
The agent loop is the core execution pattern under every framework: send the model a message plus tool definitions, let it decide whether to call a tool, execute the tool, append the result, and call the model again until it produces a final answer.
Chain-of-thought prompting asks the model to reason step by step before answering, in its simplest form by appending "Let's think step by step." Because tokens are generated sequentially, the intermediate reasoning becomes context for the final answer, effectively a scratchpad.
A circuit breaker stops an agent stuck in a loop from burning tokens: a tool errors, the agent retries, same error, twenty iterations later the user is staring at a spinner.
The context window is everything the model can see for a single request: system prompt, conversation history, retrieved documents, tool schemas, and tool results.
Few-shot prompting means showing the model three to five examples of input and desired output so it mimics the pattern rather than following abstract rules.
A Generative Adversarial Network trains two networks against each other: a Generator turns random noise into images while a Discriminator judges real from fake, each improving until the fakes pass, like a counterfeiter versus a detective.
A handoff is the swarm-style mechanism, popularized by OpenAI's Swarm framework, where an agent does not call another agent but becomes one: it returns a handoff and the framework transfers control, context, and conversation history to the target.
The LLM wrapper trap is building a thin product, model plus UI plus system prompt, that dies the moment the model provider ships your feature as a default.
Because self-attention processes all tokens simultaneously, a Transformer has no built-in sense of word order; "dog bites man" and "man bites dog" would look identical.
Prompt injection is an attack where malicious instructions are hidden in data the agent processes, such as an email or web page, and the model follows them as if they came from its operator.
ReAct, short for Reasoning plus Acting, is the loop most production agents run: the model thinks about what it needs, takes an action such as a tool call, observes the result, and repeats until it can answer.
A state space model such as Mamba processes a sequence as a signal flowing through a dynamical system rather than comparing every token to every other.
The supervisor pattern uses a central router agent that reads each request, delegates it to a specialist, reviews the output, and decides what happens next.
This is the book's central evaluation idea: the system prompt already specifies what the agent should do, so use it as the spec instead of hand-labeling test cases.
A trace is one complete agent run: the user's input, every reasoning step, retrieval, tool call, and the final answer, captured as a single top-level record.