Encyclopedia
A concise, cross-linked reference. Each entry connects to related concepts and the articles that go deeper.
Browse by topic
A
Anomaly Detection
Finding the few observations that were not produced by the process that produced the rest. The defining difficulty is not the algorithm but the base rate: at 0.5% anomalies, a detector that never fires is 99.5% accurate, and most standard metrics inherit that number rather than measuring skill.
Autocorrelation
The correlation of a series with a lagged copy of itself, measuring how long the influence of an observation persists. It is the structure that makes time-series data informative and the reason ordinary standard errors do not apply to it.
B
Bagging and Random Forests
Ensemble methods that reduce variance by averaging many models fitted to bootstrap resamples, with random forests additionally decorrelating the trees by restricting the features available at each split.
Bias-Variance Trade-off
The decomposition of a model’s expected prediction error into bias, variance, and irreducible noise, and the tension that reducing one of the first two typically increases the other.
C
Causal Graph
A drawing of assumed cause-and-effect relationships as arrows between variables, used to decide which variables must be adjusted for and which must not - a question the data alone cannot answer.
Condition Number
The ratio of the largest to the smallest curvature of a loss surface, which alone determines how fast gradient descent can converge on it.
Confounding
A variable that influences both the treatment and the outcome, so that a comparison between the treated and the untreated measures the difference between the groups as well as the effect of the treatment.
Conjugate Prior
A prior chosen so that the posterior belongs to the same family, which turns Bayesian updating into arithmetic on the parameters and makes the prior readable as a number of imagined observations.
Cross-Entropy
A measure of the difference between two probability distributions, used as the standard loss function for classification.
Cross-Validation
A resampling method that estimates a model’s test error by repeatedly fitting it on part of the data and evaluating it on the part held out.
D
E
G
H
K
k-Means Clustering
An unsupervised algorithm that partitions observations into k groups by alternately assigning points to the nearest centroid and recomputing the centroids.
k-Nearest Neighbours
A nonparametric classifier that predicts the class of a point by taking a majority vote among the k training observations closest to it.
Kullback-Leibler Divergence
The number of extra bits per symbol paid for describing one distribution with a code built for another. It is zero only when the two agree, it is never negative, and it is not symmetric, so it is a cost rather than a distance.
L
Learning Rate Schedule
A rule that changes the step size over the course of training, large early so the run can travel and small late so it can settle.
Linear Discriminant Analysis
A generative classifier that models each class as a Gaussian and inverts those models with Bayes’ theorem; assuming one covariance matrix shared by all classes gives a linear decision boundary, and one per class gives a quadratic one.
Linear Regression
A model that predicts a numeric response as a weighted sum of the predictors, fitted by minimizing squared error.
Logistic Regression
A classification model that predicts the probability of a class by passing a linear combination of predictors through the logistic function.
M
Matrix Factorisation
A model that explains a sparse table of interactions as the product of two small matrices, giving every user and every item a short vector of learned traits whose dot product predicts the missing entries.
Multiple Comparisons
The inflation of false positives that occurs whenever more than one test, metric, segment or stopping point is allowed to produce the headline. Each additional chance raises the probability that something crosses the threshold by luck alone.
Mutual Information
How many bits observing one variable tells you about another. It is zero exactly when the two are independent, it catches dependence of any shape rather than linear dependence only, and nothing computed downstream can increase it.
N
Naive Bayes
A classifier that applies Bayes’ theorem while assuming all features are conditionally independent given the class.
Neural Network
A model composed of layers of simple units, each computing a weighted sum followed by a non-linear function, fitted by gradient descent using backpropagation.
O
P
PAC Learning
A definition of learnability in which an algorithm must return, with high probability, a hypothesis whose true error is within a chosen tolerance - using a number of samples that is bounded in advance rather than discovered afterwards.
Perplexity
The exponential of a model’s average cross entropy, read as the number of equally likely options it is effectively choosing between at each step.
Precision and Recall
Two rates that split what accuracy hides: precision is the share of predicted positives that are real, and recall is the share of real positives that were found.
Pretraining and Fine-Tuning
The two-stage recipe of first training a model on a large generic corpus, then adapting it to a specific task with a much smaller labelled dataset.
Principal Component Analysis
A technique that re-expresses data in new uncorrelated coordinates ordered by how much variance each explains, allowing dimension reduction by keeping only the first few.
Q
R
Regularization
Any technique that constrains a model’s effective complexity in order to reduce variance and improve generalization, typically by penalizing large parameter values.
ROC Curve
A plot of a classifier’s true positive rate against its false positive rate as the decision threshold is swept across its whole range, summarising every available trade between the two kinds of error.
S
Softmax
A function that turns a vector of real scores into a probability distribution by exponentiating each score and dividing by the total, preserving their order while making them positive and summing to one.
Spline
A piecewise polynomial joined at chosen points called knots, constrained so that the function and its lower derivatives stay continuous there, giving local flexibility without the wild behaviour of a high-degree polynomial.
Stationarity
A property of a series whose statistical behaviour does not depend on when you look at it: the mean, the variance and the correlation structure are the same in every window. Almost every classical method assumes it, and most real series lack it.
Stochastic Gradient Descent
Gradient descent in which each step uses the gradient of a small random sample of the data rather than all of it, trading an exact direction for far more steps per unit of compute.
Support Vector Machine
A classifier that separates classes with the boundary leaving the widest possible margin, determined only by the closest training points.