Skip to content
Kudos AI

Activation Function

The non-linear function applied to a layer’s output, without which a network of any depth would collapse to a single linear transformation.

Also known as: Non-linearity, ReLU, Softmax

Understanding Activation Function

Every hidden layer applies an activation function element-wise to its linear output. The purpose is structural rather than cosmetic: with no activation, stacking layers produces another linear map and depth is wasted. The choice of function then determines how gradients behave during training, which is what makes it consequential in practice.

ReLU, which returns the input when positive and zero otherwise, is the default for hidden layers. It is trivially cheap to compute, and for positive inputs its derivative is exactly one, so it neither shrinks nor inflates gradients passing back through it. That property is what makes very deep networks trainable, and it is the activation Chollet uses throughout for hidden layers.

The older sigmoid and tanh functions squash their input into a bounded range, and there lies the problem. For large positive or negative inputs the curve flattens, so the derivative approaches zero. Since backpropagation multiplies these derivatives across layers, a few saturated units drive gradients in early layers toward nothing, and learning stalls. ReLU’s adoption is largely a response to this.

ReLU has its own failure mode: a unit whose input is always negative outputs zero and receives zero gradient, so it can never recover. Variants such as leaky ReLU, which gives a small non-zero slope for negative inputs, exist to address this. Output layers are a separate matter, chosen to match the target rather than to aid optimization: softmax for a distribution over classes, sigmoid for an independent probability, and no activation at all for unbounded regression.

How to Calculate

ReLU(x) = max(0, x); σ(x) = 1/(1 + e⁻ˣ); softmax(x)ᵢ = e^{xᵢ} / Σⱼ e^{xⱼ}

where

ReLU
rectified linear unit, the usual hidden-layer choice
σ
the sigmoid, mapping any real number into (0, 1)
softmax
normalizes a vector of scores into a probability distribution summing to 1

Frequently Asked Questions

Why did ReLU largely replace sigmoid in hidden layers?

Sigmoid saturates: for inputs far from zero its gradient is nearly zero, and backpropagation multiplies those small factors across layers, so early layers stop learning. ReLU has gradient exactly one for positive inputs, so the signal passes through undiminished.

What is the dying ReLU problem?

If a unit’s input is always negative it outputs zero and its gradient is zero, so its weights never update and the unit is permanently inert. Leaky ReLU and similar variants avoid this by giving negative inputs a small non-zero slope.

How is the output activation chosen?

By the shape of the target. Softmax for mutually exclusive multi-class problems, sigmoid for binary or independent multi-label ones, and no activation for unbounded regression. Unlike hidden activations, this choice is dictated by the task rather than by optimization.

The Bottom Line

Activation functions are what make depth worth having, and their gradient behaviour decides whether a deep network trains at all. ReLU is the sensible hidden-layer default; the output activation is fixed by the prediction task.