Part I · The Foundation

Chapter 1

How It All Started

Four unrelated events that created the AI revolution.

2,347 words · 3 diagrams · free to read

The story of AI is not a straight line. It's four separate threads — running in parallel for years, sometimes decades — that nobody thought were connected.

A gaming graphics card. A young professor labeling millions of photos. A Netflix competition. And a handful of researchers the academic establishment had mostly written off.

Each one, on its own, looked like a dead end or a curiosity. Together, they detonated.

The Four Threads That Created AI
The Four Threads That Created AI


1.1 The Long Road — AI Never Really Stopped

There's a popular narrative that AI "died" twice and was magically resurrected. It's a great story. It's also mostly wrong.

The idea of a machine that thinks is older than most people realize. In 1943, McCulloch and Pitts published the first mathematical model of an artificial neuron. In 1957, Rosenblatt built the perceptron — and the New York Times said it would one day "walk, talk, see, write, reproduce itself, and be conscious of its existence."

The hype cycle is not a new invention.

Then Minsky and Papert proved single-layer perceptrons couldn't solve the XOR problem (1969). Funding dried up — the first "AI winter." The second came in the late 1980s when expert systems turned out to be brittle and useless the moment they encountered anything unexpected.

But research never stopped. Rumelhart, Hinton, and Williams published backpropagation in 1986. Yann LeCun built the first practical convolutional neural networks in 1989. Support vector machines and random forests became workhorses in the 2000s. By the early 2010s, a practical ML ecosystem was thriving — Kaggle had turned ML into a global sport, and tools like Scikit-learn made algorithms accessible in a few lines of Python.

The people who call 2012 a "sudden breakthrough" weren't paying attention to the 20 years of work that made it possible.

AI Research Timeline: From McCulloch-Pitts (1943) to AlexNet (2012)
AI Research Timeline: From McCulloch-Pitts (1943) to AlexNet (2012)


1.2 The GPU Trick

Around the year 2000, a Stanford PhD student named Ian Buck had an insight that would eventually transform how we use computers. He wasn't thinking about AI. He was thinking about graphics cards.

Here's the problem he noticed. CPUs do things one at a time — versatile and precise, but sequential. A GPU is a factory floor — thousands of simple workers all doing the same task at once. Buck realized that matrix multiplication — the core operation behind neural networks — is embarrassingly parallel: every element can be computed independently. What took days on a CPU could take hours on a GPU.

CPU vs GPU: Why AI Needs Parallel Processing
CPU vs GPU: Why AI Needs Parallel Processing

This was an order-of-magnitude improvement, and it came from hardware being mass-produced for teenagers playing video games. NVIDIA was subsidizing the future of artificial intelligence with Call of Duty revenue, and nobody at NVIDIA knew it yet.

The catch: programming a GPU for general math in 2000 was brutal — you had to disguise your calculations as fake graphics operations. That changed in 2006 when NVIDIA released CUDA, which let programmers write normal code that ran on GPUs.

The hardware foundation was in place. It was waiting for data.


1.3 The Data Revolution

In 2007, Fei-Fei Li was an assistant professor at Princeton with a conviction that most of her colleagues thought was misguided.

Her argument was simple. The bottleneck in computer vision wasn't better algorithms. It was better data. Models were data-hungry, and nobody was feeding them.

The datasets researchers were using at the time had a few thousand images. Li wanted to build something with millions. Across thousands of categories. Labeled by hand.

She called it ImageNet.

The scale she was proposing was considered borderline absurd. This was before cloud computing was mainstream. Before crowdsourcing was a well-understood concept. Colleagues told her it was a waste of time. One senior researcher said the project would take decades to complete using graduate students.

Li had a different idea. She used Amazon Mechanical Turk — the platform that lets you pay humans small amounts to do simple tasks — to label images at scale. Workers would see an image, confirm whether it contained a golden retriever or a fire truck or a mushroom, and move on.

The engineering challenge was substantial. Quality control alone was a research problem. How do you ensure consistency across thousands of anonymous workers labeling millions of images? Li's team developed consensus mechanisms, redundant labeling, and quality filters that achieved 97%+ accuracy across the dataset.

And we still do this today. Every time you solve a CAPTCHA — "select all images with traffic lights," "click every square with a crosswalk" — you're labeling training data for self-driving car models. Fei-Fei Li's crowdsourcing idea didn't just build ImageNet. It became the foundation for how AI training data is created at scale. Millions of people, labeling images, every day, for free. They just don't know they're doing it.

By the time it was done, ImageNet contained over 14 million images across more than 21,000 categories. The competition subset — the one that would matter most — had 1.2 million training images across 1,000 categories.

Nobody had ever assembled anything like it. And in 2010, Li did something that would prove critical: she launched the ImageNet Large Scale Visual Recognition Challenge (ILSVRC). Every year, research teams would compete to build the best image classifier, measured against the same massive dataset.

The first two years were incremental. Teams used hand-engineered features — SIFT descriptors, histograms of oriented gradients, support vector machines. Error rates dropped slowly. From 28% to 26%.

The fuel was ready. It just needed a spark.


1.4 The Netflix Prize

In October 2006, Netflix offered one million dollars to anyone who could beat their recommendation algorithm by 10%. Over 40,000 teams registered. The competition ran for nearly three years.

Geoffrey Hinton's team at the University of Toronto entered with Restricted Boltzmann Machines, a type of neural network. They didn't win, but their approach was competitive with methods that had been refined for decades. That was the signal.

The winning team eventually crossed the threshold in 2009 — but their solution was so complex that Netflix never actually deployed it.

The real lesson was subtler. Neural networks — the technology the establishment had dismissed twice — could compete with decades of specialized feature engineering. Given enough data and enough compute, learning beat knowing. The pattern that would define the next fifteen years was already visible, if you knew where to look.

Most people weren't looking.


1.5 The Inflection Point — AlexNet (2012)

By 2012, the four threads were in place. GPUs were fast and programmable. ImageNet had assembled the largest labeled image dataset ever created. Neural network research had quietly continued in a handful of labs despite two decades of skepticism. And the Netflix Prize had shown that deep learning could compete.

Nobody had put them all together. Not yet.

Alex Krizhevsky was a graduate student at the University of Toronto, working under Geoffrey Hinton. His colleague Ilya Sutskever — who would later co-found OpenAI — was a PhD student in the same lab. Together, they built a deep convolutional neural network and entered the 2012 ImageNet challenge.

Their hardware budget was not impressive. They had two NVIDIA GTX 580 gaming GPUs. Not a supercomputer. Not a government-funded cluster. Two graphics cards you could buy at Best Buy for $500 each.

They called their network AlexNet.

It had eight layers. 60 million parameters. 650,000 neurons. By today's standards, it was tiny — GPT-4 is estimated at over a trillion parameters. But in 2012, it was the deepest network anyone had successfully trained on a dataset this large.

Krizhevsky and his team made several technical choices that broke with convention. They used ReLU activation functions instead of the standard sigmoid — a change that sounds minor but solved the vanishing gradient problem that had plagued deep networks for years. They used dropout regularization to prevent overfitting — randomly disabling neurons during training so the network couldn't memorize the data. They used data augmentation — flipping, cropping, and adjusting images to artificially expand the training set.

And they trained the whole thing on GPUs. Not CPUs. Two gaming GPUs running in parallel, splitting the network across both cards because neither had enough memory to hold the full model.

The results came in on September 30, 2012.

AlexNet achieved a top-5 error rate of 15.3%. The second-place entry, using traditional hand-engineered features, scored 26.2%.

Read those numbers again. AlexNet didn't win by a percentage point or two. It cut the error rate nearly in half. In a field where annual progress was measured in fractions of a percent, this was an earthquake.

The gap between first and second place was larger than all the progress made in the previous three years of the competition combined.

Jensen Huang, looking back years later, put it simply: "The invention of deep learning began with two flagship Fermi GPUs."

All four threads converged in this one entry. GPU compute made it possible to train. ImageNet provided the data to learn from. A deep architecture with the right technical innovations — ReLU, dropout, data augmentation — extracted features no human engineer had designed. And the right people — Hinton, who had spent decades defending neural networks through two AI winters, and his brilliant students — knew how to put it all together.

AlexNet wasn't a single cause. It was a confluence. Better optimization, better regularization, better hardware, better data, and a decade of quiet progress in unsupervised pretraining and representation learning all contributed. But AlexNet was the proof point — the result so dramatic that the field couldn't ignore it.

Before AlexNet, deep learning was a niche interest pursued by a stubborn minority. After it, deep learning became the dominant paradigm in machine learning research. Within two years, nearly every competitive entry in ImageNet used deep neural networks.

The major results that followed — speech recognition, machine translation, game playing, protein folding, language generation — all built on this convergence of data, compute, and architecture. Two GPUs. One dataset. Eight layers. And the vindication of an idea that had been rejected twice.


1.6 Google Brain and the Cat

While AlexNet was upending computer vision, a parallel effort inside Google was about to reshape the entire tech industry.

In 2011, Andrew Ng and Jeff Dean — two of the most respected names in machine learning and systems engineering, respectively — co-founded Google Brain as a research project within Google X, the company's moonshot division.

Their first major experiment was audacious. They connected 16,000 CPU cores across 1,000 machines into a single massive neural network — 1 billion connections in total. Then they fed it 10 million unlabeled thumbnails randomly selected from YouTube videos.

No labels. No categories. No instructions. Just raw data and compute.

After training for three days, they examined what the network had learned. One set of neurons had specialized. It had developed a detector for something very specific.

Faces.

Another set of neurons responded strongly to a different pattern.

Cats.

"The system had discovered the concept of a cat itself," the team reported. "No one had ever told it what a cat was."

The result became famous. It was covered by the New York Times, the BBC, every tech blog on the internet. "Google's AI learned to recognize cats" became shorthand for a profound insight: given enough data and enough compute, neural networks could discover structure in the world without being taught. Unsupervised learning at scale actually worked.

The Google Brain cat paper used CPUs, not GPUs — 16,000 cores across 1,000 machines. It was comically inefficient by modern standards. Within two years, researchers would replicate similar results on a single GPU workstation. But the proof of concept sent shockwaves through the industry.

Google's leadership noticed. Jeff Dean kept building.

In March 2013, Google acquired DNNResearch Inc. — Geoffrey Hinton's startup, which was essentially just Hinton and two former students, including Alex Krizhevsky. The acquisition was the result of a bidding war. Baidu, the Chinese search giant, was the other serious bidder. Hinton reportedly set up an auction that lasted four days.

Google won. Hinton joined Google Brain, bringing with him not just his expertise but the symbolic weight of the researcher who had kept neural networks alive through two AI winters.

And then Google did something unusual.

In November 2015, Google open-sourced TensorFlow — the internal machine learning framework that powered Google Brain. The strategy was elegant: if everyone trains on your framework, talent flows to you. But the effect was broader — TensorFlow made deep learning accessible to anyone with a laptop and a GPU.

The arms race accelerated.

Facebook hired Yann LeCun. Baidu poached Andrew Ng. Microsoft, Amazon, and Apple all launched AI labs. The arms race for talent mirrored the arms race for compute.

The pattern from AlexNet was being internalized across the industry: neural networks plus data plus compute equals breakthroughs. The logical next step was to throw more of everything at the problem.

The foundation had been laid. What came next — the transformer, GPT, ChatGPT, the race to AGI — was all built on top of these four threads converging in the early 2010s.


Key Takeaways

  • Four threads converged: GPU compute, ImageNet's data, stubborn neural network researchers, and the Netflix Prize's signal that deep learning could compete
  • AlexNet cut image recognition errors nearly in half with two gaming GPUs — not a single cause, but a confluence of better hardware, better data, better optimization (ReLU, dropout), and decades of quiet research
  • Google Brain's unsupervised cat detector proved the concept at scale, triggering a talent war that launched the modern AI boom
  • The talent war mattered as much as the technology — Google Brain, DeepMind, Baidu, and Facebook all recruited aggressively, and the release of TensorFlow democratized deep learning for everyone

Paper Spotlight: "ImageNet Classification with Deep Convolutional Neural Networks" (Krizhevsky, Sutskever, Hinton, 2012). 60 million parameters. 650,000 neurons. ReLU activation. Dropout regularization. Trained on those two GPUs. The paper that proved deep learning works at scale. Over 150,000 citations and counting. If you read one paper from the deep learning revolution, make it this one.

Keep reading

Chapter 2 is free too.

Kindle, paperback and hardcover on Amazon.com, Kindle on Amazon.in. 22 chapters, 65,000 words, 80-plus original diagrams like the ones above.

Buy on Amazon.com
Read chapter 2: Before the Transformer →