Definition

Positional Encoding

Because self-attention processes all tokens simultaneously, a Transformer has no built-in sense of word order; "dog bites man" and "man bites dog" would look identical. Positional encoding fixes this by adding a unique mathematical fingerprint, built from sine and cosine waves at different frequencies, to each token's embedding before it enters the model. Position is supplied by math, not learned.

Positional Encoding diagram from Intelligence at Scale
Diagram from chapter 3, Attention Is All You Need

Explained in

Chapter 3: Attention Is All You Need Free

The paper behind every major AI model today. Google invented it. Then watched someone else ship it.

Related terms

This is one term. The chapter is the argument.

Intelligence at Scale: 22 chapters, 65,000 words, 80-plus diagrams. Kindle, paperback and hardcover on Amazon.

Buy on Amazon.com
← All terms