Optimization for Learning
Every model on this site is fitted by the same loop: look at the slope, take a step. What decides whether that loop converges in forty steps or diverges in three is not the model - it is curvature, noise, and the size of the step. All three are measurable before the first epoch runs.
Sign in to take quizzes, earn XP, and unlock stages as you reach 90% mastery.
The Step, and the Edge of Stability
25 min · 100 XPWhat a gradient step is minimising, why a minibatch gradient is the full gradient plus noise rather than a different gradient, and the exact step size above which descent stops descending.
Open lesson →Sign in to take the 3-question quiz.
Curvature, and What Momentum Buys
25 min · 100 XPThe condition number is the single number that predicts how slow plain descent will be, and momentum changes that prediction from κ to √κ - a claim precise enough to check to six decimal places.
Open lesson →Sign in to take the 3-question quiz.
Noise, Schedules, and Adaptive Methods
25 min · 100 XPStochastic gradient descent with a fixed step does not converge - it settles into a ball whose radius grows as √η. Decaying the step removes the floor; Adam rescales each coordinate; and on a long enough run the fastest method is the one that finishes last.
Open lesson →Sign in to take the 3-question quiz.