Skip to content
Kudos AI
Lire en français
Neural Networks

Backpropagation and Gradient Descent

How a neural network learns: the loss as a function of weights, gradient descent, and backpropagation as the chain rule applied backwards, with every partial derivative of a small network computed by hand and checked against autograd.

8 min readKudos AI

Prerequisites: Linear Regression from First Principles, Logistic Regression and Classification

The loss drawn as a surface with the gradient pointing uphill and the step going the other way, then the chain rule walked backwards through the article's worked network until the loss falls from 0.66125 to 0.2457.

A neural network is a chain of simple operations, each with adjustable weights. Training means finding weights that make the output right. Two ideas do all the work: gradient descent, which says which way to move the weights, and backpropagation, which computes the gradient efficiently no matter how deep the chain is.

Backpropagation is often described as difficult. It is the chain rule, applied in a specific order. This article computes every derivative of a small network by hand and then checks the results against an automatic differentiation engine.

A. Layers, weights, and the loss

A dense layer applies a linear transformation followed by an element-wise non-linearity:

a=g(Wx+b),\mathbf{a} = g(W\mathbf{x} + \mathbf{b}) ,

where WW is a weight matrix, b\mathbf{b} a bias vector, and gg an activation function. Without gg, stacking layers would be pointless: a composition of linear maps is just another linear map, and the network could represent nothing a single layer could not.

We use the ReLU, g(z)=max⁡(0,z)g(z) = \max(0, z), whose derivative is 11 for z>0z > 0 and 00 for z<0z < 0.

A loss function scores the output against the target. For a scalar output we use

L=12(y^−y)2,L = \tfrac{1}{2}(\hat y - y)^2 ,

the 12\tfrac12 being a convenience that cancels when differentiated.

The key reframing: with the data held fixed, LL is a function of the weights. Training is minimising that function.

B. Gradient descent

The gradient ∇L\nabla L points in the direction of steepest increase, so to decrease LL we step against it:

w←w−η∂L∂w,w \leftarrow w - \eta \frac{\partial L}{\partial w} ,

with η\eta the learning rate. Too small and training crawls; too large and it overshoots - the same failure demonstrated numerically in Logistic Regression, where a step of 0.50.5 made the objective substantially worse.

So the entire problem reduces to computing ∂L/∂w\partial L / \partial w for every weight. A network can have millions, and computing each independently would be hopeless. Backpropagation gets all of them in a single backward pass.

C. The network we will differentiate

Two inputs, a two-unit ReLU hidden layer, one linear output. Biases are zero.

x=[12],W(1)=[0.10.30.20.4],W(2)=[0.5−0.5],y=1.\mathbf{x} = \begin{bmatrix}1\\2\end{bmatrix},\quad W^{(1)} = \begin{bmatrix}0.1 & 0.3\\ 0.2 & 0.4\end{bmatrix},\quad W^{(2)} = \begin{bmatrix}0.5 & -0.5\end{bmatrix},\quad y = 1 .

Forward pass.

z(1)=W(1)x=[0.1(1)+0.3(2)0.2(1)+0.4(2)]=[0.71.0].\mathbf{z}^{(1)} = W^{(1)}\mathbf{x} = \begin{bmatrix}0.1(1) + 0.3(2)\\ 0.2(1) + 0.4(2)\end{bmatrix} = \begin{bmatrix}0.7\\ 1.0\end{bmatrix} .

Both entries are positive, so ReLU passes them through unchanged:

a(1)=[0.71.0].\mathbf{a}^{(1)} = \begin{bmatrix}0.7\\ 1.0\end{bmatrix} . y^=z(2)=0.5(0.7)+(−0.5)(1.0)=0.35−0.5=−0.15.\hat y = z^{(2)} = 0.5(0.7) + (-0.5)(1.0) = 0.35 - 0.5 = -0.15 . L=12(−0.15−1)2=12(−1.15)2=12(1.3225)=0.66125.L = \tfrac12(-0.15 - 1)^2 = \tfrac12(-1.15)^2 = \tfrac12(1.3225) = 0.66125 .

D. The backward pass

We now walk backwards, carrying the derivative of LL with respect to each quantity as we go. Each step is one application of the chain rule.

Step 1 - the output.

∂L∂y^=y^−y=−0.15−1=−1.15.\frac{\partial L}{\partial \hat y} = \hat y - y = -0.15 - 1 = -1.15 .

Step 2 - the output weights. Since y^=W(2)a(1)\hat y = W^{(2)}\mathbf{a}^{(1)}, we have ∂y^/∂Wj(2)=aj(1)\partial \hat y / \partial W^{(2)}_j = a^{(1)}_j, so

∂L∂W(2)=∂L∂y^ a(1)=−1.15[0.71.0]=[−0.805−1.15].\frac{\partial L}{\partial W^{(2)}} = \frac{\partial L}{\partial \hat y}\,\mathbf{a}^{(1)} = -1.15 \begin{bmatrix}0.7 & 1.0\end{bmatrix} = \begin{bmatrix}-0.805 & -1.15\end{bmatrix} .

Note the structure: the gradient of a weight is the error signal arriving at it multiplied by the activation flowing into it. That pattern holds at every layer.

Step 3 - back through the output layer. To continue we need how LL changes with the hidden activations:

∂L∂a(1)=∂L∂y^ W(2)=−1.15[0.5−0.5]=[−0.5750.575].\frac{\partial L}{\partial \mathbf{a}^{(1)}} = \frac{\partial L}{\partial \hat y}\,W^{(2)} = -1.15\begin{bmatrix}0.5 & -0.5\end{bmatrix} = \begin{bmatrix}-0.575 & 0.575\end{bmatrix} .

The error is distributed backwards in proportion to the weights that carried the signal forward.

Step 4 - through the ReLU. Multiply element-wise by the activation's derivative. Both pre-activations were positive, so both derivatives are 11:

∂L∂z(1)=[−0.5750.575].\frac{\partial L}{\partial \mathbf{z}^{(1)}} = \begin{bmatrix}-0.575 & 0.575\end{bmatrix} .

This is where dead units come from. Had a pre-activation been negative, the ReLU derivative would be 00 and the gradient would be annihilated - no error signal reaches any weight feeding that unit, and it cannot learn. A unit stuck negative for all inputs is a dead unit, and it is the standard motivation for variants such as leaky ReLU.

Step 5 - the input weights. Same pattern as step 2: error signal times incoming activation, here the input itself.

∂L∂W(1)=∂L∂z(1)x⊤=[−0.5750.575][12]=[−0.575−1.150.5751.15].\frac{\partial L}{\partial W^{(1)}} = \frac{\partial L}{\partial \mathbf{z}^{(1)}} \mathbf{x}^{\top} = \begin{bmatrix}-0.575\\ 0.575\end{bmatrix}\begin{bmatrix}1 & 2\end{bmatrix} = \begin{bmatrix}-0.575 & -1.15\\ 0.575 & 1.15\end{bmatrix} .

Every gradient in the network, from one forward pass and one backward pass.

E. Checking against autograd

Hand derivations are error-prone, and there is no reason to trust one that has not been checked:

Python

Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.

Both print True: the hand-computed gradients agree with PyTorch's autograd to floating-point precision.

F. Taking the step

Every weight moves against its own gradient. With η=0.1\eta = 0.1:

W(2)←[0.5−0.5]−0.1[−0.805−1.15]=[0.5805−0.385],W^{(2)} \leftarrow \begin{bmatrix}0.5 & -0.5\end{bmatrix} - 0.1\begin{bmatrix}-0.805 & -1.15\end{bmatrix} = \begin{bmatrix}0.5805 & -0.385\end{bmatrix} , W(1)←[0.10.30.20.4]−0.1[−0.575−1.150.5751.15]=[0.15750.4150.14250.285].W^{(1)} \leftarrow \begin{bmatrix}0.1 & 0.3\\ 0.2 & 0.4\end{bmatrix} - 0.1\begin{bmatrix}-0.575 & -1.15\\ 0.575 & 1.15\end{bmatrix} = \begin{bmatrix}0.1575 & 0.415\\ 0.1425 & 0.285\end{bmatrix} .

Recomputing the forward pass with both updated matrices gives y^=0.2989\hat y = 0.2989 and

L=12(0.2989−1)2=0.2457,L = \tfrac12(0.2989 - 1)^2 = 0.2457 ,

down from 0.661250.66125. One step cut the loss by more than half. Note that y^\hat y moved from −0.15-0.15 towards the target y=1y = 1 - it overshot none of the way, it simply moved in the right direction, which is all a gradient step promises.

The figure below is this network with every number on it: the activations on the units, and on each weight the gradient that reaches it. Check any of them against the arithmetic above. Then move the learning rate. The lesson is right that the steps shrink on their own near a minimum, and it is worth knowing how close the cliff is on the other side: at 0.5, five times the rate used here, the first oversized step drives one hidden unit negative and the third drives the other, ReLU then blocks every gradient beneath them, and the network sits at a loss of 0.5 for ever. That is not divergence. It is death, and it takes three steps.

Interactive: one backward pass, every number on the table

The lesson’s network. Take a step, then try a larger learning rate.

Forward values, and the gradient on each weight

-0.5750.575-1.1501.150-0.805-1.150x1 = 1x2 = 2h10.700h21.000y-hat-0.150
Prediction
-0.1500
Loss
0.6612
dL/dy-hat
-1.1500
Steps taken
0

The prediction is -0.15 against a target of 1, so the loss is 0.6612 and the first derivative is -1.15. Read that sign as a direction rather than a verdict: negative means the loss falls as the prediction rises, which is exactly right for a prediction sitting below its target. Every gradient on the diagram is that one number pushed backwards through the weights, which is the whole trick - computed once, reused for every parameter that feeds the output.

G. Batches, and why the algorithm scales

Real training does not use one example. Stochastic gradient descent computes the gradient on a small random mini-batch and steps, repeating over the dataset. The batch gradient is noisier than the full-dataset gradient but vastly cheaper, and the noise is often helpful - it can knock the parameters out of poor regions.

Backpropagation's efficiency is what makes any of this feasible: one forward pass and one backward pass yield the derivative with respect to every parameter, with cost proportional to the forward pass rather than to the number of parameters. Computing each of a million partial derivatives separately by finite differences would need a million forward passes.

The gradient is local, and the loss surface is not convex. Unlike least squares, a network's loss has many minima and saddle points, and gradient descent offers no guarantee of finding the best one. In practice good-enough minima are common, but "it converged" is not "it found the optimum".

Key takeaways

  • Training reframes the loss as a function of the weights and minimises it by gradient descent, w←w−η ∂L/∂ww \leftarrow w - \eta\,\partial L/\partial w.
  • Non-linear activations are essential; stacked linear layers collapse to one.
  • Backpropagation is the chain rule applied backwards, reusing each intermediate derivative.
  • Every weight gradient is incoming activation × outgoing error signal - the same rule at every layer.
  • Our worked network: L=0.66125L = 0.66125, ∂L/∂W(2)=[−0.805,−1.15]\partial L/\partial W^{(2)} = [-0.805, -1.15], ∂L/∂W(1)=[[−0.575,−1.15],[0.575,1.15]]\partial L/\partial W^{(1)} = [[-0.575, -1.15], [0.575, 1.15]], confirmed against autograd.
  • ReLU units with negative pre-activations pass zero gradient and can die.

What's next

The same backward pass trains the architecture behind modern language models, but those models need a mechanism for letting each position in a sequence consult every other. That mechanism is Attention and Self-Attention.

References & further reading

  • François Chollet, Deep Learning with Python, Manning (2nd edition, MEAP), 2020· Kudos AI reference library

Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.

Related reading

10 min readNeural Networks

What Actually Makes Training Converge

A two per cent change in the learning rate separates a converged run from one five orders of magnitude away, a condition number predicts the convergence rate to six decimal places, and stochastic gradient descent with a fixed step never converges at all - it settles into a ball whose radius grows as the square root of the step. Every figure here was computed on a problem whose exact optimum is known.

OptimizationDeep LearningMachine Learning
5 min readNeural Networks

The Slowest Direction Sets the Pace

The step size you are allowed is fixed by the steepest direction and the number of steps you need is fixed by the flattest, so the cost of gradient descent is their ratio. The same least-squares fit, to the same ten decimal places, takes 1742 steps in one basis, 147 in a rescaled one and exactly 1 in an orthonormal one, and momentum buys back the square root of the ratio rather than the ratio.

Machine LearningMathematics
8 min readNeural Networks

What Is a Neural Network?

Layers as parameterised transformations, the forward pass, and why depth and non-linearity are not optional: a proof that no single linear layer can compute XOR, and a two-layer network that does, worked entirely by hand.

Deep LearningMathematics
← Back to all articles