Skip to content
Kudos AI

Convolutional Neural Network

A neural network that applies learned filters across an input’s spatial extent, sharing weights so the same pattern is detected wherever it occurs.

Also known as: CNN, ConvNet

One 3x3 kernel swept across an image with the same nine weights lighting up at every stop - weight sharing as a claim about the world, not a saving.

Understanding Convolutional Neural Network

A fully connected layer applied to an image treats every pixel as an unrelated input, which throws away the fact that nearby pixels are related and that a pattern means the same thing wherever it appears. It is also ruinously expensive: connecting a modest image to a modest hidden layer requires millions of weights.

A convolutional layer replaces this with a small filter, perhaps 3×3, slid across the whole input. At each position it computes a weighted sum of the local patch, producing a feature map that records where in the image that filter’s pattern was found. Crucially, the same weights are used at every position.

Weight sharing gives two benefits at once. The parameter count collapses, since a 3×3 filter has nine weights regardless of image size, so a layer costs a few thousand parameters rather than millions. And because the same filter is applied everywhere, a pattern learned in one part of the image is automatically detected elsewhere, rather than having to be learned separately for each location.

Depth then builds a hierarchy. Early layers, seeing only small patches, learn simple local structure such as edges and colour transitions. Later layers operate on the feature maps produced by earlier ones, so their effective view of the original image is larger, and they combine simple features into parts and eventually whole objects. Chollet notes that convolutional networks and backpropagation were both well understood by 1990, and that what changed after 2012 was the availability of data and compute rather than the core ideas.

Example of Convolutional Neural Network

A 3×3 filter with large positive weights down one column and large negative weights down another responds strongly wherever brightness changes horizontally, and weakly on flat regions. Slid across the image it produces a map of vertical edges. Filters like this are not designed; they emerge from training.

Consider the parameter arithmetic on a 224×224 colour image. A fully connected layer to 1,000 units needs 224 × 224 × 3 × 1,000 ≈ 150 million weights. A convolutional layer with 64 filters of size 3×3×3 needs 64 × 27 ≈ 1,728 weights plus biases, roughly five orders of magnitude fewer.

Pooling layers are typically interleaved to reduce spatial resolution, which widens the region of the original image each later unit responds to and adds tolerance to small shifts. The overall pattern of the network is a gradual trade of spatial detail for semantic abstraction.

Frequently Asked Questions

Why are convolutional layers better than fully connected ones for images?

They exploit two facts a fully connected layer ignores: that relevant structure is local, and that a pattern means the same thing wherever it appears. This gives far fewer parameters and much better generalization from the same amount of data.

What does pooling do?

It downsamples a feature map, typically by taking the maximum over each small window. This reduces computation, enlarges the region of the input that later layers effectively see, and makes the representation somewhat insensitive to small translations.

Are convolutional networks only for images?

No. The same idea applies wherever data has a grid structure with meaningful locality: one-dimensional convolutions are used for audio and time series, and convolutional architectures have been applied to text, though attention-based models now dominate there.

The Bottom Line

Convolutional networks bake locality and translation reuse directly into the architecture, so a pattern is learned once and detected everywhere at a fraction of the parameter cost. Stacking them yields a learned hierarchy from edges to objects, which is why they defined computer vision for a decade.