Part I · The Foundation

Chapter 2

Before the Transformer

What models existed before Transformers — and why they weren't enough.

1,161 words · 3 diagrams · free to read

Every revolution has a "before." A period where smart people built impressive things that almost worked. Where you could see the shape of the future, but couldn't quite reach it.

The story of AI language models before the Transformer is exactly that. A series of brilliant innovations, each one solving part of the puzzle. Each one hitting the same wall.

To understand why the Transformer mattered so much, you need to understand what came before it. Not because the old architectures were bad — they weren't. They powered Google Translate, Siri's early voice recognition, and the first wave of machine translation that actually felt usable.

But they all shared the same fundamental flaw. And that flaw made scaling impossible.


2.1 Teaching Machines to Read

Before neural networks could generate language, they had to learn what words mean.

The idea started earlier than most people realize. In 2003, Yoshua Bengio published "A Neural Probabilistic Language Model" — the first paper to show that neural networks could learn word representations from raw text. The insight was profound. The timing was terrible. GPUs weren't ready. Datasets were tiny. The idea sat dormant for a decade.

Then Tomas Mikolov at Google made it fast enough to be useful. In 2013, he published Word2Vec — a shallow neural network trained to predict a word from its context. The results were wild:

king - man + woman ≈ queen
Paris - France + Italy ≈ Rome
bigger - big + small ≈ smaller

Nobody programmed this. It emerged from co-occurrence patterns in billions of words. Word2Vec proved that neural networks could learn meaningful representations of language from raw data. Every model that followed built on this.

Word Embeddings: Meaning as Geometry in Vector Space
Word Embeddings: Meaning as Geometry in Vector Space


2.2 RNNs, LSTMs, and the Sequential Bottleneck

The dominant architecture for language before Transformers was the Recurrent Neural Network (RNN). The idea: language is sequential, so process words one at a time, left to right, carrying forward what you've learned.

Think of it like reading a book while trying to remember the entire plot. No notes. No going back. Just your memory — getting hazier with every page.

The problem: vanishing gradients. Error signals get multiplied by fractions at each step. Do that 100 times and the signal is effectively zero. The network can't learn relationships between words that are far apart.

LSTMs (Hochreiter & Schmidhuber, 1997) added memory gates — three of them. An input gate controls what new information gets written. A forget gate decides what to discard. An output gate determines what gets passed forward. These gates let the network selectively remember across long sequences. Google used LSTMs for Translate for years. But they were still sequential — to process word 100, you had to wait for words 1 through 99. GPUs are massively parallel, but you can't parallelize an LSTM.

The Sequential Bottleneck: RNNs vs Transformers
The Sequential Bottleneck: RNNs vs Transformers


Then in 2014, Ilya Sutskever, Oriol Vinyals, and Quoc Le at Google introduced Sequence-to-Sequence — the encoder-decoder architecture. One LSTM reads the entire input and compresses it into a single vector. A second LSTM takes that vector and generates the output, word by word.

Like summarizing a novel into a single tweet — then asking someone else to reconstruct the story from that tweet alone. It worked shockingly well for translation. But the bottleneck was obvious. Everything the model knew about the input had to fit in one fixed-size vector. Short sentences? Fine. Long paragraphs? Information got crushed.


2.3 The Attention Mechanism

In 2014, Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio fixed the bottleneck.

Their insight: instead of forcing the decoder to work from a single compressed vector, let it look back at any part of the input while generating each output word. They called this attention.

When translating "Le chat est sur la table" to English, the decoder generates "cat" by focusing on "chat." It generates "table" by focusing on "la table." Each output word chooses which input words matter most.

No more information bottleneck. Long sentences stopped losing information. Translations got significantly better across every benchmark.

But even with attention, the architecture was still an RNN underneath. Still sequential. Still one word at a time. Attention solved the information problem but not the speed problem.


2.4 Google Neural Machine Translation (2016)

In November 2016, Google switched Google Translate from a phrase-based statistical system to GNMT — an 8-layer LSTM encoder-decoder with attention. Translation errors dropped 55-85%. Choppy, literal translations suddenly felt natural.

The real insight came when Google trained a single model on all languages simultaneously. Instead of degrading, the model got better at all languages. It could even translate between language pairs it had never seen during training — zero-shot translation. The model had learned a representation of language itself.

This was the first clear signal: scale changes everything. Capabilities emerged at scale that simply didn't exist at smaller sizes. It was also the last major system built on the LSTM architecture.


2.5 Why This Mattered — And Why It Wasn't Enough

Let's take stock of where things stood by late 2016.

Evolution of Sequence Models
Evolution of Sequence Models

Every innovation in this chapter solved a real problem and hit the same wall: sequential processing. Word 50 had to wait for words 1 through 49. The hardware was ready — GPUs could handle thousands of operations in parallel. The architectures weren't.

The field needed something that threw away the sequential constraint entirely.


Key Takeaways

  • Bengio's 2003 paper showed neural networks could learn word representations — the idea sat dormant for a decade until Word2Vec made it fast enough to use
  • Word2Vec proved neural networks could learn word meaning from raw text — king minus man plus woman equals queen, with nobody programming it
  • RNNs and LSTMs introduced sequential processing with memory gates (input, forget, output), but both hit the same wall: one word at a time, wasting GPU parallelism
  • Seq2Seq gave us encoder-decoder — compress the input into a vector, generate the output from it. Brilliant, but everything had to fit through a single bottleneck
  • Attention removed the information bottleneck by letting models focus on relevant input words, but the architecture was still sequential underneath
  • Google's GNMT achieved zero-shot translation between language pairs it was never trained on — an early sign that scale produces emergent capabilities
  • Every pre-Transformer innovation solved a real problem and hit the same limit: sequential processing couldn't scale

Paper Spotlight: "Neural Machine Translation by Jointly Learning to Align and Translate" (Bahdanau, Cho, Bengio, 2015). The paper that introduced the attention mechanism. Instead of forcing an entire sentence through a single vector bottleneck, it let the model selectively focus on relevant parts of the input while generating each output word. This single concept — attention — became the entire foundation of the Transformer architecture two years later. When Vaswani et al. titled their paper "Attention Is All You Need," they weren't being cute. They were being literal.

That was the Transformer. Next chapter.

Keep reading

Chapter 3 is free too.

Kindle, paperback and hardcover on Amazon.com, Kindle on Amazon.in. 22 chapters, 65,000 words, 80-plus original diagrams like the ones above.

Buy on Amazon.com
Read chapter 3: Attention Is All You Need →