Part I · The Foundation
Chapter 3
Attention Is All You Need
The paper behind every major AI model today. Google invented it. Then watched someone else ship it.
2,327 words · 4 diagrams · free to read
3.1 Google's Translation Problem
Google's GNMT had achieved zero-shot translation across 100+ languages (Chapter 2). But the second discovery was even more important.
Bigger models performed better. Not just a little better. Predictably, reliably, consistently better. Double the parameters, and translation quality went up. Double them again, it went up again. There was no obvious ceiling.
This violated the intuition that most machine learning researchers had at the time. The standard assumption was that at some point, adding more parameters would lead to overfitting — the model would memorize training data instead of learning general patterns. You'd hit diminishing returns.
That's not what happened. More parameters, more data, more compute — the relationship was almost embarrassingly straightforward. More is more.
This "scaling" pattern would define the next decade of AI research. It would fuel a compute arms race that consumed billions of dollars. It would reorder the priorities of every major tech company on the planet.
3.2 "Attention Is All You Need"
June 2017. Eight researchers at Google publish a paper with a title so casual it almost sounds like a dare.
"Attention Is All You Need."
The authors: Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser, and Illia Polosukhin. In a move that broke academic convention, all eight listed themselves as "equal contributors" with randomized author order. No first author. No hierarchy. Eight people, equal credit, one paper.
LSTMs were the state of the art, but they processed sequences one token at a time — a bottleneck that limited parallelization and made long-range dependencies hard to capture.
Vaswani and his co-authors proposed something radical.
Throw away recurrence entirely. No sequential processing. No hidden state carried forward word by word. Instead, use only one mechanism: attention.
The idea of attention wasn't new. Bahdanau had introduced it in 2014 as an add-on to RNNs — a way for the model to look back at relevant parts of the input while generating output. But it was always a supplement. The RNN was still doing the heavy lifting.
The Transformer paper asked: what if attention is the only thing you need?
Here's how self-attention works, stripped of the math.
Imagine you're reading a sentence: "The cat sat on the mat because it was tired."
What does "it" refer to? The cat or the mat? You know it's the cat because you understand the semantic relationship between "tired" and "cat." Mats don't get tired.
In self-attention, every word in the sentence can attend to every other word simultaneously. The word "it" doesn't have to wait for the model to sequentially process "The," "cat," "sat," "on," "the," "mat," "because" before it can figure out what it refers to. It looks at all of them at once. It computes a relevance score with every other word in the sentence, weights them, and produces a new representation that captures the relationships.
Every word does this. Simultaneously. In parallel.
The model computes three vectors for each word: a Query (what am I looking for?), a Key (what do I contain?), and a Value (what information do I carry?). Attention scores are calculated by comparing each word's Query against every other word's Key. High scores mean high relevance. Those scores are used to create a weighted sum of Values. The result is a representation of each word that's informed by every other word in the sentence.
It's like every word in the sentence having a conversation with every other word, all at the same time.

But the Transformer didn't stop there. It introduced multi-head attention — running multiple attention patterns in parallel.
Think of it this way. When you read a sentence, you're tracking multiple types of relationships simultaneously. Syntactic structure: which words are the subject and object? Semantic meaning: what concepts are related? Coreference: what does "it" refer to? Temporal order: what happened first?
Multi-head attention handles this by running, say, 8 or 12 separate attention mechanisms in parallel. Each "head" learns to focus on a different type of relationship. One head might learn syntax. Another might learn coreference. Another might learn semantic similarity. The model figures out the division of labor on its own during training.
The outputs from all heads are concatenated and projected back into a single representation. The result is richer than any single attention pattern could produce.
There's one more piece. Without recurrence, the model has no sense of word order — "Dog bites man" and "Man bites dog" look identical. The solution: positional encoding. Each position in the sequence gets a unique mathematical fingerprint — generated using sine and cosine functions at different frequencies — added to the word embeddings before they enter the model. The model doesn't learn position. It's told position, through math. An elegant hack that solves a fundamental problem.

Put it all together — self-attention, multi-head attention, positional encoding, feed-forward layers, residual connections — and you get the full Transformer architecture.

Why did the Transformer win? Four reasons.
Parallelism. Every word is processed simultaneously. No more sequential bottleneck. A sentence with 500 words takes roughly the same wall-clock time as a sentence with 50 words, given enough compute. This meant Transformers could exploit the full power of modern GPUs. Training times dropped dramatically.
Scalability. Because the architecture is inherently parallel, it scales almost linearly with hardware. Add more GPUs, get proportionally faster training. This is exactly what you want if your strategy is "make the model bigger."
Long-range dependencies. Self-attention connects every word to every other word directly. There's no information bottleneck. Word 1 is exactly as accessible as word 499. The vanishing gradient problem that plagued RNNs? Gone.
Universality. The same architecture works for any sequence task. Translation. Summarization. Question answering. Code generation. It doesn't care what language you're working in. It doesn't care if the input is text, code, music, or protein sequences. It's a general-purpose sequence processor.
Every major AI model you've heard of is built on this architecture. ChatGPT. Claude. Gemini. Llama. DALL-E. Whisper. Stable Diffusion. Sora. Every single one. The Transformer is to modern AI what the transistor is to modern computing.
And here's the part that should make every startup founder pay attention: all eight authors eventually left Google. Noam Shazeer cofounded Character AI. Aidan Gomez cofounded Cohere. Llion Jones cofounded Sakana AI. Jakob Uszkoreit founded Inceptive. Ashish Vaswani and Niki Parmar cofounded Essential AI. Illia Polosukhin cofounded NEAR Protocol. Łukasz Kaiser went to OpenAI. Eight people wrote the most important AI paper of the century, and Google lost every single one of them. The company that invented the Transformer couldn't retain the people who built it.
Eight researchers. One paper. Over 140,000 citations and counting. The world hasn't been the same since.
3.3 BERT — Google's Gift (That Google Didn't Use)
October 2018. Just over a year after the Transformer paper, Jacob Devlin and his colleagues at Google AI released BERT — Bidirectional Encoder Representations from Transformers. If the Transformer was the engine, BERT was the first car built around it that ordinary people could drive.
BERT introduced two key innovations. The first was bidirectional context. Previous language models read text left to right. BERT read in both directions simultaneously. When you read "I went to the bank to deposit my check," you understand "bank" means a financial institution because of "deposit" that comes after it. A left-to-right model doesn't have that luxury. BERT does.
The second was Masked Language Modeling. During training, BERT randomly hides 15% of the words in a sentence and trains the model to predict them from surrounding context — a massive game of fill-in-the-blank, played over billions of sentences.
The numbers: BERT Base had 110 million parameters across 12 layers. BERT Large had 340 million across 24 layers. Trained on Wikipedia (2.5 billion words) and BookCorpus (800 million words). That's it. Two datasets. No proprietary data. No secret sauce.
And here's what made BERT revolutionary: you could actually use it. BERT could be fine-tuned on a single GPU in a few hours. Inference ran on a CPU. A laptop. A startup with no GPU budget could download it from GitHub and have a state-of-the-art language model in production by the end of the week.
Before BERT, every NLP task was its own bespoke engineering project. BERT changed that. Its representations were contextual — the vector for "bank" in "deposit my check" is completely different from "bank" in "sat on the river bank." Same word. Different meaning. Different numbers. These context-aware embeddings became a universal foundation: take BERT's output and feed it into any downstream model.
This is what "pre-train then fine-tune" means in practice. BERT learned the structure of language once, for everyone. You just plugged in your specific task on top. The paradigm had arrived, and it would redefine the entire field.
The results were absurd. BERT didn't just beat the state of the art on one benchmark. It crushed eleven simultaneously. Question answering. Sentiment analysis. Named entity recognition. Textual entailment. Every NLP task that researchers had spent years optimizing for — BERT swept all of them with a single model, fine-tuned for a few hours each. The paper has been cited over 100,000 times.
Within months, the entire field pivoted. If you submitted an NLP paper without a BERT baseline, reviewers sent it back. Research groups abandoned years of specialized work overnight. Then came Multilingual BERT — a single model trained on 104 languages. One architecture. One training run. Every language.
BERT didn't just set a new bar. It made the old bar irrelevant.
3.4 The Decoder-Only Insight
The original Transformer had two halves. An encoder that reads the input and a decoder that generates the output. BERT used only the encoder. It could understand text brilliantly — classify it, analyze it, answer questions about it — but it couldn't generate new text from scratch. BERT was a reader, not a writer.
OpenAI made the opposite bet. GPT used only the decoder half. It read text left to right and predicted the next token. Always the next token. Nothing else.
This seems like a limitation. BERT can see the whole sentence at once. GPT can only see what came before the current position. How is that better?
It's better because generation is harder and more general than understanding. If you can generate coherent text, you can implicitly understand it too. A model that can write a convincing essay about quantum physics understands quantum physics — or at least enough about it to be useful. But a model that understands quantum physics can't necessarily write about it.
The decoder-only architecture also has a property that turned out to be transformative: it scales gracefully. The training objective is dead simple — predict the next token. You don't need labeled data. You don't need paired inputs and outputs. You just need text. Lots and lots of text. The entire internet is a training dataset.
BERT needed task-specific fine-tuning. You had to define your task, prepare labeled examples, train a classification head, and evaluate. GPT just needed a prompt. You asked it a question in natural language, and it answered. No fine-tuning. No special training head. Just text in, text out.
This is the key insight that the entire LLM revolution is built on. The decoder-only Transformer, trained to predict the next token on massive amounts of text, produces a general-purpose reasoning engine. It's not specialized for any task. It's not optimized for any domain. It just predicts tokens. And that turns out to be enough.
Nobody fully understood this in 2018. Not even OpenAI. But OpenAI scaled from 117 million parameters (GPT-1) to 175 billion (GPT-3) in just two years — we'll trace that journey in the next chapter. Each jump in scale revealed new capabilities: translation, code generation, mathematical reasoning, creative writing. Abilities that nobody programmed, nobody trained for, nobody expected.
The researchers called these emergent abilities — capabilities that appear at scale without being explicitly trained. The decoder-only Transformer was a box that kept producing new surprises every time you made it bigger.
Google had the architecture. They chose the encoder. OpenAI chose the decoder. That choice — encoder versus decoder, understanding versus generation — is why we say "GPT" and not "BERT" when we talk about the AI revolution.

One paper. Eight authors. Two architectural choices. The entire trajectory of modern AI — from Google's BERT to OpenAI's GPT — traces back to this fork in the road.
Google invented the Transformer. Google released BERT. They published the paper, open-sourced the code, posted the weights on GitHub. And then they went back to doing research.
Meanwhile, a small nonprofit lab in San Francisco was paying very close attention.
Key Takeaways
- Self-attention lets every word attend to every other word simultaneously — parallel, not sequential. That's what made Transformers scalable
- BERT proved that pre-training on unlabeled text, then fine-tuning for specific tasks, democratized NLP — suddenly small teams could build state-of-the-art models
- The decoder bet won because predicting the next token needs no labeled data — the entire internet becomes your training set
- Emergent abilities — translation, code, reasoning — appeared at scale without being explicitly trained
- All eight Transformer authors left Google. The company that invented the architecture couldn't keep the people who built it
Paper Spotlight: "Attention Is All You Need" (Vaswani, Shazeer, Parmar, Uszkoreit, Jones, Gomez, Kaiser, Polosukhin, 2017). Arguably the most important paper in modern AI. Introduced the Transformer architecture: self-attention allows every token to attend to every other token simultaneously, multi-head attention runs multiple attention patterns in parallel, positional encoding handles word order without sequential processing. The architecture behind every major AI model today: GPT, Claude, Gemini, Llama, DALL-E, Whisper, Stable Diffusion, Sora. The paper was 15 pages. Its impact is incalculable.
And then, on November 30, 2022, OpenAI shipped the fastest-growing consumer application in history.
Keep reading
The remaining 19 chapters are in the book.
Kindle, paperback and hardcover on Amazon.com, Kindle on Amazon.in. 22 chapters, 65,000 words, 80-plus original diagrams like the ones above.