Understanding Backpropagation
A neural network is a deeply nested composition of functions: each layer takes the previous layer’s output, applies a linear transformation and a non-linearity, and passes the result on. To train it with gradient descent you need the partial derivative of the final loss with respect to every individual weight, including weights buried many layers deep. Backpropagation is the procedure that computes them.
The mechanism is the chain rule, applied systematically. A forward pass runs the input through the network and records each layer’s intermediate values. The backward pass then starts at the loss and works toward the input: at each layer it receives the derivative of the loss with respect to that layer’s output, and uses it to compute two things, the derivative with respect to that layer’s weights (which is what the optimizer needs) and the derivative with respect to that layer’s input (which is passed back to the layer before).
Efficiency is the entire reason the algorithm matters. Estimating each of a million partial derivatives separately, by nudging one weight at a time and re-running the network, would need a million forward passes. Backpropagation reuses the shared intermediate quantities, obtaining every gradient in a single backward sweep whose cost is comparable to the forward pass. Without that reduction, training networks of modern size would be arithmetically impossible.
The attribution is genuinely layered. The 1986 work of Rumelhart, Hinton and Williams is what brought the method to the field’s attention and set off the connectionist revival, but the underlying technique, reverse-mode automatic differentiation applied to layered networks, had been derived independently by several earlier authors. It is fairer to describe 1986 as the moment backpropagation became widely known than as the moment it was invented.
How to Calculate
∂L/∂w⁽ˡ⁾ = δ⁽ˡ⁾ · (a⁽ˡ⁻¹⁾)ᵀ, where δ⁽ˡ⁾ = (W⁽ˡ⁺¹⁾)ᵀ δ⁽ˡ⁺¹⁾ ⊙ σ′(z⁽ˡ⁾)
where
- L
- the loss being minimized
- w⁽ˡ⁾, W⁽ˡ⁾
- the weights of layer l
- a⁽ˡ⁻¹⁾
- the activation output by the previous layer
- z⁽ˡ⁾
- the pre-activation input to layer l
- δ⁽ˡ⁾
- the error signal at layer l, propagated backwards
- σ′
- the derivative of the activation function
- ⊙
- element-wise (Hadamard) multiplication
Example of Backpropagation
Consider a trivial two-layer chain where the loss depends on b, b depends on a, and a depends on the weight w. The chain rule gives ∂L/∂w = (∂L/∂b)(∂b/∂a)(∂a/∂w): the gradient is a product of local derivatives along the path.
Backpropagation computes this right-to-left. It first evaluates ∂L/∂b at the output, then multiplies by ∂b/∂a to obtain ∂L/∂a, then multiplies by ∂a/∂w to obtain the gradient for w. Each intermediate result is reused rather than recomputed, and in a real network with many paths converging on a node, the contributions arriving from all downstream paths are summed.
This ordering is also the source of the vanishing-gradient problem. Because the gradient is a product of many local derivatives, if those factors are consistently smaller than one, the product shrinks exponentially with depth, and early layers receive almost no signal. Activation functions such as ReLU, and architectural devices such as residual connections, exist largely to keep that product well-behaved.
Frequently Asked Questions
Is backpropagation the same thing as gradient descent?
No, and the distinction matters. Backpropagation computes the gradients; gradient descent decides what to do with them. You could feed backpropagation’s output to any gradient-based optimizer, such as Adam or momentum, and it would still be backpropagation supplying the derivatives.
Why does backpropagation need the values from the forward pass?
The local derivative at each layer generally depends on the values that flowed through it. That is why frameworks retain the intermediate activations during the forward pass, and why memory use grows with network depth and batch size during training but not during inference.
What is the vanishing gradient problem?
Because backpropagation multiplies local derivatives together across layers, small factors compound. In deep networks with saturating activations this drives gradients in the early layers toward zero, so those layers barely learn. ReLU activations, careful initialization, normalization layers, and residual connections are the standard mitigations.
The Bottom Line
Backpropagation is the reason deep learning is computationally feasible: it turns the problem of finding millions of partial derivatives into a single organized backward sweep of the chain rule. Every gradient-based training procedure in modern deep learning depends on it.