Encyclopedia
A concise, cross-linked reference. Each entry connects to related concepts and the articles that go deeper.
Browse by topic
A
Anomaly Detection
Finding the few observations that were not produced by the process that produced the rest. The defining difficulty is not the algorithm but the base rate: at 0.5% anomalies, a detector that never fires is 99.5% accurate, and most standard metrics inherit that number rather than measuring skill.
Autocorrelation
The correlation of a series with a lagged copy of itself, measuring how long the influence of an observation persists. It is the structure that makes time-series data informative and the reason ordinary standard errors do not apply to it.
B
Bagging and Random Forests
Ensemble methods that reduce variance by averaging many models fitted to bootstrap resamples, with random forests additionally decorrelating the trees by restricting the features available at each split.
Bayes’ Theorem
A rule for updating the probability of a hypothesis in light of new evidence, by inverting a conditional probability.
Bias-Variance Trade-off
The decomposition of a model’s expected prediction error into bias, variance, and irreducible noise, and the tension that reducing one of the first two typically increases the other.
C
Causal Graph
A drawing of assumed cause-and-effect relationships as arrows between variables, used to decide which variables must be adjusted for and which must not - a question the data alone cannot answer.
Confidence Interval
A range computed from data by a procedure that, repeated over many samples, contains the true value a stated proportion of the time. The stated proportion is a property of the procedure, not of any particular interval it produces.
Confounding
A variable that influences both the treatment and the outcome, so that a comparison between the treated and the untreated measures the difference between the groups as well as the effect of the treatment.
Cross-Validation
A resampling method that estimates a model’s test error by repeatedly fitting it on part of the data and evaluating it on the part held out.
E
H
K
k-Means Clustering
An unsupervised algorithm that partitions observations into k groups by alternately assigning points to the nearest centroid and recomputing the centroids.
k-Nearest Neighbours
A nonparametric classifier that predicts the class of a point by taking a majority vote among the k training observations closest to it.
L
Linear Discriminant Analysis
A generative classifier that models each class as a Gaussian and inverts those models with Bayes’ theorem; assuming one covariance matrix shared by all classes gives a linear decision boundary, and one per class gives a quadratic one.
Linear Regression
A model that predicts a numeric response as a weighted sum of the predictors, fitted by minimizing squared error.
Logistic Regression
A classification model that predicts the probability of a class by passing a linear combination of predictors through the logistic function.
M
Matrix Factorisation
A model that explains a sparse table of interactions as the product of two small matrices, giving every user and every item a short vector of learned traits whose dot product predicts the missing entries.
Maximum Likelihood Estimation
A method of fitting a model by choosing the parameter values that make the observed data most probable.
Multiple Comparisons
The inflation of false positives that occurs whenever more than one test, metric, segment or stopping point is allowed to produce the headline. Each additional chance raises the probability that something crosses the threshold by luck alone.
O
P
p-value
The probability of observing data at least as extreme as the data in hand, computed under the assumption that the null hypothesis is true. It measures how unusual the sample would be in a world where the effect is absent, and nothing else.
Precision and Recall
Two rates that split what accuracy hides: precision is the share of predicted positives that are real, and recall is the share of real positives that were found.
Principal Component Analysis
A technique that re-expresses data in new uncorrelated coordinates ordered by how much variance each explains, allowing dimension reduction by keeping only the first few.
R
Regularization
Any technique that constrains a model’s effective complexity in order to reduce variance and improve generalization, typically by penalizing large parameter values.
ROC Curve
A plot of a classifier’s true positive rate against its false positive rate as the decision threshold is swept across its whole range, summarising every available trade between the two kinds of error.
S
Spline
A piecewise polynomial joined at chosen points called knots, constrained so that the function and its lower derivatives stay continuous there, giving local flexibility without the wild behaviour of a high-degree polynomial.
Stationarity
A property of a series whose statistical behaviour does not depend on when you look at it: the mean, the variance and the correlation structure are the same in every window. Almost every classical method assumes it, and most real series lack it.
Statistical Power
The probability that a test rejects the null hypothesis when a specified alternative is true. It is the chance of finding an effect that is genuinely there, and it is fixed by the design before any data are collected.