Skip to content
Kudos AI
Lire en français
Neural Networks

What Is a Neural Network?

Layers as parameterised transformations, the forward pass, and why depth and non-linearity are not optional: a proof that no single linear layer can compute XOR, and a two-layer network that does, worked entirely by hand.

8 min readKudos AI

Prerequisites: Linear Regression from First Principles

Three stacked linear layers collapsing on screen into a single one, and then failing to collapse once a nonlinearity is put between them.

A neural network is not a model of the brain. It is a chain of simple, adjustable transformations, stacked so that each one operates on the output of the last. Chollet is blunt about the terminology: the name is a reference to neurobiology, and he says two names that could equally have been chosen are layered representations learning and hierarchical representations learning. The "deep" in deep learning is not a claim about depth of understanding - it stands for the number of layers stacked on top of each other.

This article builds the object from the bottom up: what one layer is, why a non-linearity between layers is mathematically necessary rather than decorative, and what depth actually buys. The central argument is settled with a proof and a worked example small enough to check on paper.

A. A layer is a parameterised transformation

A dense layer takes a vector in, multiplies it by a matrix, adds a bias, and applies a function element-wise:

a=g(Wx+b).\mathbf{a} = g(W\mathbf{x} + \mathbf{b}) .

Three pieces:

  • WW and b\mathbf{b} are the layer's weights - the numbers that get adjusted during training. Everything a network knows lives here.
  • gg is the activation function, applied to each component independently.
  • x\mathbf{x} is whatever the previous layer produced.

A network is a stack of these. Chollet's framing is that each layer learns a representation of the data, and depth means successive layers of increasingly useful representations, all learned automatically from exposure to training data rather than hand-designed. This is the contrast with what he calls shallow learning, which extracts one or two layers of representation - say, a pixel histogram followed by a classification rule - and requires a human to decide what those features should be.

B. Why the activation function is not optional

Suppose we drop gg and stack two purely linear layers:

h=W(1)x,y=W(2)h=W(2)(W(1)x)=(W(2)W(1))x.\mathbf{h} = W^{(1)}\mathbf{x}, \qquad \mathbf{y} = W^{(2)}\mathbf{h} = W^{(2)}\left(W^{(1)}\mathbf{x}\right) = \left(W^{(2)}W^{(1)}\right)\mathbf{x} .

Matrix multiplication is associative, so the two matrices collapse into one. Concretely, with

W(1)=[2003],W(2)=[1101],W^{(1)} = \begin{bmatrix}2 & 0\\ 0 & 3\end{bmatrix},\qquad W^{(2)} = \begin{bmatrix}1 & 1\\ 0 & 1\end{bmatrix},

the composition is

W(2)W(1)=[2303],W^{(2)}W^{(1)} = \begin{bmatrix}2 & 3\\ 0 & 3\end{bmatrix},

a single layer. Add a hundred more linear layers and it is still a single layer. Depth without non-linearity buys nothing at all - not a slightly weaker model, but exactly the same set of representable functions.

The most common activation is the ReLU, g(z)=max⁡(0,z)g(z) = \max(0, z), which passes positive values through unchanged and clamps negatives to zero. It is barely more than a kink, and that kink is enough.

The figure uses a different two-layer network, without biases. At its starting input both units are on and the network equals its collapsed linear map; drag the first input or second input until a unit switches off, and ReLU parts company with the collapse.

Interactive: the collapse, and what stops it

Take the nonlinearity away and the slice becomes a straight line.

Pre-activations
0.700, 1.000
Output
-0.150
Collapsed layer says
-0.150
Units switched off
0
Distance from affine
0.150

The two layers multiply out to the single row [-0.050, -0.050], so without a nonlinearity three layers or thirty express exactly what one expresses: the distance from affine is 2e-16, which is zero to machine precision. Put ReLU back and it is 0.150. But look at where you are. Both units are on, so the rectifier is doing nothing here, and the network returns -0.150 while its collapsed form returns -0.150: the same number. That is also true at the lesson’s own input of (1, 2). Depth does not buy a curve. It buys a set of REGIONS, each of which is still an affine map, and the kinks in the slice are where a pre-activation crosses zero. Move the inputs across one and the two answers part company. One counting note while the layer is in view: a dense layer from four inputs to three outputs holds 15 parameters, not 12. The bias vector is the part that gets forgotten.

C. A function no single layer can compute

To show that the kink matters, we need a task a linear model provably cannot do. XOR is the standard one: two binary inputs, output 11 when they differ.

x1x_1x2x_2XOR
000
011
101
110

Claim. No function of the form y=w1x1+w2x2+by = w_1x_1 + w_2x_2 + b reproduces this table.

Proof. Suppose one did. Take the rows in turn:

  • (0,0)↦0(0,0) \mapsto 0 forces b=0b = 0.
  • (1,0)↦1(1,0) \mapsto 1 forces w1+b=1w_1 + b = 1, so w1=1w_1 = 1.
  • (0,1)↦1(0,1) \mapsto 1 forces w2+b=1w_2 + b = 1, so w2=1w_2 = 1.
  • (1,1)↦0(1,1) \mapsto 0 forces w1+w2+b=0w_1 + w_2 + b = 0.

Substituting the first three into the fourth gives 1+1+0=21 + 1 + 0 = 2, and we needed 00. The contradiction is unavoidable, so no such w1,w2,bw_1, w_2, b exist. ■\blacksquare

Geometrically, a linear model can only carve the input space with a single straight boundary, and the two XOR classes sit on opposite diagonals - no line separates them.

D. Two layers, worked by hand

Now add one hidden layer of two ReLU units. The weights below are chosen, not trained, so that every number can be checked:

W(1)=[1111],b(1)=[0−1],W(2)=[1−2],b(2)=0.W^{(1)} = \begin{bmatrix}1 & 1\\ 1 & 1\end{bmatrix},\quad \mathbf{b}^{(1)} = \begin{bmatrix}0\\ -1\end{bmatrix},\quad W^{(2)} = \begin{bmatrix}1 & -2\end{bmatrix},\quad b^{(2)} = 0 .

Both hidden units compute x1+x2x_1 + x_2; they differ only in their bias, so the first fires whenever the sum is positive and the second only once the sum reaches 22. The output layer subtracts twice the second from the first.

Take x=(1,1)\mathbf{x} = (1,1) step by step:

z(1)=W(1)x+b(1)=[1+1+01+1−1]=[21],\mathbf{z}^{(1)} = W^{(1)}\mathbf{x} + \mathbf{b}^{(1)} = \begin{bmatrix}1+1+0\\ 1+1-1\end{bmatrix} = \begin{bmatrix}2\\ 1\end{bmatrix}, a(1)=max⁡(0,z(1))=[21],y=1(2)+(−2)(1)=0.✓\mathbf{a}^{(1)} = \max(0, \mathbf{z}^{(1)}) = \begin{bmatrix}2\\ 1\end{bmatrix}, \qquad y = 1(2) + (-2)(1) = 0 . \checkmark

All four inputs:

x\mathbf{x}z(1)\mathbf{z}^{(1)}a(1)\mathbf{a}^{(1)}yyXOR
(0,0)(0,0)(0,−1)(0, -1)(0,0)(0, 0)000
(0,1)(0,1)(1,0)(1, 0)(1,0)(1, 0)111
(1,0)(1,0)(1,0)(1, 0)(1,0)(1, 0)111
(1,1)(1,1)(2,1)(2, 1)(2,1)(2, 1)000

Exact on all four rows.

Watch row one. The second unit's pre-activation is −1-1, and the ReLU clamps it to 00. That clamping is the entire source of the network's extra power: it is the one place where the composition stops being linear. Remove it - let a(1)=z(1)\mathbf{a}^{(1)} = \mathbf{z}^{(1)} - and section B's collapse applies again, so the network would once more be unable to compute XOR.

Python

Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.

E. What depth actually buys

Chollet gives the clearest picture of this. Imagine a red sheet of paper and a blue one, stacked and then crumpled together into a ball. The ball is your input data; each sheet is a class. The model's job is to find a transformation that uncrumples the ball so the two sheets are cleanly separable again.

A deep network does this incrementally: rather than searching for one enormously complicated transformation, it decomposes the job into a long chain of elementary ones, each a movement your fingers could make on the paper ball. Our XOR network is the smallest possible instance - the hidden layer folds the input space so that the two classes, previously inseparable by any line, become separable by one.

This is why depth is a genuinely different resource from width. Both add parameters, but depth adds composition, and composition is what turns a chain of simple transformations into a complicated one.

Capacity is not understanding. Chollet is careful to add that such a model is essentially a very high-dimensional curve fitted by gradient descent, with enough parameters that it could fit almost anything - train it long enough and it will end up memorising. The ability to represent a function is not the ability to generalise from data, which is the concern of the bias-variance tradeoff.

F. From representation to learning

We chose the XOR weights by hand. Real networks discover them, and the mechanism is a loss function - what Chollet also calls the objective function - that takes the network's prediction and the true target and computes a distance score capturing how badly the network did. Training adjusts the weights to shrink that score.

Getting from "here is a score" to "here is how each of a million weights should change" is the subject of the next article.

Key takeaways

  • A layer is a parameterised transformation g(Wx+b)g(W\mathbf{x} + \mathbf{b}); the weights are everything the network knows.
  • Stacked linear layers collapse into a single linear layer - W(2)W(1)W^{(2)}W^{(1)} is just another matrix - so a non-linearity between them is mathematically necessary, not a tuning choice.
  • No single linear layer can compute XOR; the four constraints force 1+1+0=01+1+0 = 0, a contradiction.
  • One hidden layer of two ReLU units computes XOR exactly, and the only non-linear step in it is a single clamp of −1-1 to 00.
  • Depth supplies composition: a complicated transformation decomposed into a chain of elementary ones - Chollet's uncrumpling of a folded data manifold.
  • Enough capacity to represent a function is not the same as generalising from data.

What's next

We set the XOR weights by hand. Finding them automatically means computing how the loss responds to every weight in the network at once, which is exactly what Backpropagation and Gradient Descent does.

References & further reading

  • François Chollet, Deep Learning with Python, Manning (2nd edition, MEAP), 2020· Kudos AI reference library

Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.

Related reading

8 min readNeural Networks

Backpropagation and Gradient Descent

How a neural network learns: the loss as a function of weights, gradient descent, and backpropagation as the chain rule applied backwards, with every partial derivative of a small network computed by hand and checked against autograd.

Deep LearningOptimizationMathematics
8 min readBuilding a Language Model

Tokenization and Embeddings

How text becomes numbers a model can train on: building a vocabulary, why byte pair encoding never needs an unknown token, the embedding layer as a lookup that is provably one-hot times a matrix, and why position has to be added back in by hand.

Generative AINatural Language ProcessingDeep Learning
7 min readProbability Foundations

Probability from Zero: The Language of Uncertainty

Build probability from the ground up: possible worlds, the sample space, the two basic axioms, and the addition and multiplication rules, each derived rather than asserted, with worked numeric examples.

ProbabilityMathematicsArtificial Intelligence
← Back to all articles