Don't tell the computer what to look for. Give it 1.2 million labeled photos and let it learn the features itself.
This page walks through the whole paper: the idea, the network, the tricks that made it train, the results, and what has held up since.
ImageNet asks a model to label a photo with one of 1,000 classes, from mushrooms to container ships. Before 2012, the best systems used features that people designed, then fed them to a simple classifier. They had hit a wall around 45–47% top-1 error.
Only the last step learns. If the handmade features miss something, the classifier can never see it.
Every layer is trained together from the labels. The network decides what features matter.
Five convolutional layers find patterns. Three fully connected layers combine them into a decision. Pick a layer to see what it does. Block height shows the spatial size and width shows the number of channels.
The convolution layers do almost all the arithmetic but hold only 4% of the weights. The fully connected layers hold 96% of the weights. That is why the paper needs dropout on the FC layers, and why later networks such as GoogLeNet and ResNet cut those layers down.
None of these was brand new in 2012. The contribution was showing that together they make a network this large trainable, without it simply memorizing the training set.
tanh flattens out for large inputs, so its gradient shrinks toward zero and learning stalls. ReLU has a slope of exactly 1 for any positive input. A 4-layer net on CIFAR-10 reached 25% training error six times faster.
Each training step switches off a random half of the neurons, so no neuron can rely on a particular partner. At test time all neurons are on and outputs are halved. Without it the paper reports substantial overfitting. It costs about twice as many iterations.
Random crops and mirror flips cost nothing and give the model many slightly different views. A second trick adds small color shifts along the main color directions of ImageNet (PCA), teaching the model that lighting changes don't change the object. That alone cut top-1 error by over 1%.
One 3 GB card couldn't hold the network, so each GPU got half the filters. They only talk to each other at layer 3 and in the FC layers (solid crossing lines). The paper compares against a one-GPU net with half the filters, which is smaller, so the 1.7% gain is not a like-for-like test.
A strongly active filter dampens its neighbors at the same pixel, borrowed from biology. Reported −1.4% top-1. Batch normalization replaced it in 2015.
3×3 pooling windows placed 2 pixels apart, so they overlap. Reported −0.4% top-1. Small enough that it could be run-to-run noise.
The learning rate was divided by 10 by hand whenever validation error stopped improving. Bias set to 1 gives ReLUs positive inputs early so they start learning. The paper notes weight decay lowered training error too, not only test error.
Top-1 error means the first guess is wrong. Top-5 error means the right label is not in the five best guesses. Lower is better. Hover a bar for details.
What this proves: an end-to-end CNN beats hand-built pipelines at this scale, by a margin far larger than noise. What it doesn't prove: that each individual trick is needed. The headline 15.3% also uses 7 models, 10 crops per test image, and extra pre-training data, so a single model is closer to 18%.
Taking out any one of the middle convolution layers raised top-1 error by about 2%. That was early evidence that depth itself carries the performance. The paper shows that it matters but does not explain why.
Each image produces 4,096 numbers in the last hidden layer. Images that are close in this space show the same kind of object, even when their pixels look very different (paper Figure 4). The bars above are an illustration. The same idea powers embedding search today.
| Claim | Evidence in the paper | Since then |
|---|---|---|
| ReLU trains much faster | 6× faster to 25% error, small net on CIFAR-10 | Held up |
| Dropout prevents overfitting | Without it, "substantial overfitting" | Held up |
| Depth is important | Removing a middle layer costs ≈2% top-1 | Held up VGG, ResNet |
| Augmentation reduces overfitting | PCA color shift −1% top-1; crops needed to avoid overfitting | Held up |
| Response normalization helps | −1.4% top-1, single comparison | Faded BatchNorm |
| Overlapping pooling helps | −0.4% top-1, no error bars | Faded |
| Two-GPU split helps | −1.7% vs a smaller one-GPU net | Faded hardware workaround |