Understanding Learning Rate Schedule
The step size is the one hyperparameter that has to be wrong at some point during training if it is held constant. Early on the parameters are far from any minimum and large steps are what makes progress; late on the gradient estimate is dominated by mini-batch noise, and a large step turns that noise into permanent error. A schedule resolves the conflict by making the step a function of time.
The requirement is precise. Robbins and Monro showed that a stochastic approximation converges when the step sizes satisfy two conditions at once: their sum diverges, so that the iterate can still reach an optimum however far away it starts, and the sum of their squares is finite, so that the injected noise is summable. A constant step meets the first and fails the second, which is exactly why it stalls at a floor. A step decaying like 1/t² meets the second and fails the first, and can run out of travel before arriving.
Warmup - raising the rate from near zero over the first few hundred or few thousand steps - is often described in the same breath as decay, but it answers a different question. The stability limit on a gradient step is η < 2/L, and in an untrained network L moves quickly in the first steps. Starting under the limit and rising into it avoids a divergence that is not about noise at all.
In practice a schedule is paired with an adaptive method rather than chosen instead of one. Adam and momentum change the direction and effective size of the step; the schedule governs the overall scale. The two are solving the two halves of the same problem - fast early progress, and a closing noise ball at the end - which is why almost every published training recipe specifies both.
How to Calculate
Σ_t η_t = ∞, Σ_t η_t² < ∞
where
- η_t
- the learning rate at step t
- Σ_t η_t = ∞
- the run retains enough total travel to reach the optimum from anywhere
- Σ_t η_t² < ∞
- the noise injected over the whole run is finite, so the iterate can settle
Example of Learning Rate Schedule
A least-squares problem run with mini-batches of 8 and a fixed step of 0.20 settles at a root-mean-square distance of 0.109343 from the optimum and stays there, however long it runs. Halving the step to 0.10 moves the floor only to 0.075358, because the radius scales as √η.
Replacing the constant with η_t = 0.20/(1 + t/500) - the same starting rate, the same data, the same seed - brings the run to 0.009909 over the same horizon, eleven times closer, and it is still improving when stopped.
How fast to decay is itself a choice: over this horizon τ = 500 reaches 0.009909, τ = 2000 reaches 0.019304 and τ = 10,000 is still at 0.040247. A schedule that decays too slowly spends the run at a floor it is trying to leave.
Advantages and Disadvantages
Pros
- Removes the noise floor that pins a constant-step run short of the optimum.
- Allows a large early step for fast progress without paying for it at the end.
- Costs nothing: it is a change to one scalar per step, not to the model or the data.
Cons
- Adds its own hyperparameters - the initial rate, the decay shape, the horizon - which interact with batch size.
- A schedule tuned for one run length can be badly wrong for another, since most shapes are defined relative to the total number of steps.
- Decaying too early freezes progress at whatever the parameters happen to be.
Frequently Asked Questions
Why not just stop training when the loss plateaus?
Because a plateau under a constant step is usually not a minimum - it is the edge of a noise ball whose radius is set by the step size. Decaying the rate at that moment typically produces a further drop in loss, which is what makes step-decay schedules look so dramatic on a training curve.
Does Adam make a schedule unnecessary?
No. Adam rescales the step per coordinate, which is conditioning, not noise reduction - it still settles into a ball, and by moving further per step it settles into a slightly wider one. On long runs a plain SGD with a schedule routinely ends closer to the optimum than a constant-step Adam.
Is cosine decay better than step decay?
The difference between reasonable schedules is small compared with the difference between having one and not. Cosine is popular because it needs only a total step budget and has no cliff; step decay is easier to reason about when the training length may change mid-run.
The Bottom Line
A learning-rate schedule is what turns stochastic gradient descent from a method that circles the optimum into one that reaches it. The two Robbins-Monro conditions say what any workable schedule must do, and the choice of shape matters far less than obeying them.