Skip to content
Kudos AI

Noise, Schedules, and Adaptive Methods

Stochastic gradient descent with a fixed step does not converge - it settles into a ball whose radius grows as √η. Decaying the step removes the floor; Adam rescales each coordinate; and on a long enough run the fastest method is the one that finishes last.

IntermediateModule 325 min · 100 XP
An SGD trajectory spiralling into a visible cloud around the optimum, the cloud shrinking by only 30% as the step halves, then closing to a point once the step decays - and a three-way race where the early leader finishes last.

This is a premium lesson

Sign in and enrol to read the full lesson, run the code, take the quiz, and earn XP toward the path badge.