Encyclopedia
A concise, cross-linked reference. Each entry connects to related concepts and the articles that go deeper.
Browse by topic
A
B
Backpropagation
The algorithm that computes the gradient of a neural network’s loss with respect to every weight, by applying the chain rule backwards through the network.
Bellman Equation
The self-consistency condition that the utility of a state equals its immediate reward plus the discounted value of the best action available from it, averaged over the outcomes that action cannot control.
C
Condition Number
The ratio of the largest to the smallest curvature of a loss surface, which alone determines how fast gradient descent can converge on it.
Conjugate Prior
A prior chosen so that the posterior belongs to the same family, which turns Bayesian updating into arithmetic on the parameters and makes the prior readable as a number of imagined observations.
G
K
Kalman Filter
The exact filtering algorithm for a continuous state that moves linearly with Gaussian noise and is measured linearly with Gaussian noise, carrying the whole belief as a mean and a variance.
Kullback-Leibler Divergence
The number of extra bits per symbol paid for describing one distribution with a code built for another. It is zero only when the two agree, it is never negative, and it is not symmetric, so it is a cost rather than a distance.
M
P
PAC Learning
A definition of learnability in which an algorithm must return, with high probability, a hypothesis whose true error is within a chosen tolerance - using a number of samples that is bounded in advance rather than discovered afterwards.
Principal Component Analysis
A technique that re-expresses data in new uncorrelated coordinates ordered by how much variance each explains, allowing dimension reduction by keeping only the first few.
S
Softmax
A function that turns a vector of real scores into a probability distribution by exponentiating each score and dividing by the total, preserving their order while making them positive and summing to one.
Spline
A piecewise polynomial joined at chosen points called knots, constrained so that the function and its lower derivatives stay continuous there, giving local flexibility without the wild behaviour of a high-degree polynomial.
Stochastic Gradient Descent
Gradient descent in which each step uses the gradient of a small random sample of the data rather than all of it, trading an exact direction for far more steps per unit of compute.
Support Vector Machine
A classifier that separates classes with the boundary leaving the widest possible margin, determined only by the closest training points.