Understanding Softmax
A network computing class or token scores produces arbitrary real numbers, which cannot be read as probabilities: they may be negative and they do not sum to anything in particular. Softmax is the standard repair. Exponentiating makes every entry positive and preserves order, and dividing by the sum makes the entries a distribution. Nothing else about the scores is changed, so the largest score remains the most probable outcome.
The exponential is not an arbitrary choice of positive function. It is what makes the output a distribution in the exponential family, and it is what makes the gradient of cross entropy with respect to the scores reduce to the predicted distribution minus the observed one - a subtraction, with no division and no chain of factors. That simplicity is why the pair is used together so consistently.
Two properties matter in practice. Softmax is shift-invariant: adding a constant to every score changes nothing, which says the scores are meaningful only relative to each other. Implementations exploit this by subtracting the largest score before exponentiating, because exp of a large number overflows while exp of a non-positive number cannot. And dividing every score by a temperature sharpens the distribution as the temperature falls toward zero and flattens it toward uniform as the temperature grows.
The name is misleading in a way worth naming. Softmax is a smooth stand-in for the argmax, not for the maximum: it returns a distribution over positions, which becomes a one-hot vector at the winning position as the temperature goes to zero.
How to Calculate
\mathrm{softmax}(\mathbf{z})_i = \frac{e^{z_i}}{\sum_{j=1}^{K} e^{z_j}}
where
- z_i
- the raw score, or logit, for outcome i
- K
- the number of outcomes, such as classes or vocabulary tokens
Example of Softmax
Take the scores 2, 1 and 0. Exponentiating gives 7.389056, 2.718282 and 1, summing to 11.107338, so the distribution is 0.665241, 0.244728 and 0.090031. A gap of one unit in the scores has become a ratio of about 2.72 in the probabilities, which is the exponential at work.
Add 100 to every score and the answer is unchanged to the last digit, because the added factor cancels between numerator and denominator. The naive computation would not survive it: exp(1000) overflows a double outright, which is exactly why implementations subtract the largest score first.
Changing the temperature moves the same three scores a long way. At a temperature of 0.5 the distribution is 0.8668, 0.1173 and 0.0159; at 1 it is 0.6652, 0.2447, 0.0900; at 10 it is 0.3672, 0.3322 and 0.3006, which is nearly uniform over three outcomes.
Frequently Asked Questions
Why not simply divide each score by the total?
Because scores can be negative, which would produce negative probabilities, and a zero total would divide by zero. Exponentiating first removes both problems, and it is also what gives cross entropy its simple gradient.
Does a high softmax probability mean the model is confident?
It means the winning score was far above the others, which is not the same thing. A model can produce a sharply peaked distribution on an input quite unlike anything it was trained on, and the peak carries no information about that.
The Bottom Line
Softmax turns scores into a distribution by exponentiating and normalising. It preserves order, ignores any constant added to every score, sharpens or flattens with temperature, and is the smooth stand-in for argmax rather than for max.