Skip to content
Kudos AI

Statistics

Estimation, inference, and the discipline of saying how far a number can be trusted. The foundations under every model that claims to have learned something.

75 items

Learning paths (12)

Probability and Statistical Foundations

Reason about uncertainty precisely, then meet the central problem of learning from data: separating the error you can remove from the error you cannot.

Supervised Machine Learning

Derive the workhorse supervised methods rather than merely calling them: least squares, logistic regression, shrinkage penalties, and tree ensembles.

Unsupervised Learning

Find structure in data that has no response to predict, and face the consequence squarely: with no y there is no held-out error, so every choice you make has to be defended some other way.

Moving Beyond Linearity

Keep least squares and change what you regress on: fixed basis functions buy curvature, constraints buy smoothness, and a penalty buys a curve that chooses its own flexibility.

Learning Probabilistic Models

When the data are complete, learning a probability model is counting - the derivative of the log likelihood does the rest. When variables are hidden there is nothing to count, and the repair is to guess the counts, refit, and repeat until the likelihood stops rising.

Classification Methods Compared

There is one classifier no method can beat, and it needs the answer to build. Everything else - nearest neighbours, discriminant analysis, logistic regression - is a different guess at what it would have done, and the guesses fail in different directions.

Statistical Inference

What a sample can and cannot tell you about the population behind it: how an estimator misses, what a confidence interval actually promises, and what a p-value is - together with the three places where each of those is routinely read as something stronger than it is.

Causal Inference

Why a comparison between the treated and the untreated can carry the wrong sign, what randomisation actually buys, and the rule that says which variables to adjust for - including the ones that make the answer worse.

Time Series

What breaks when observations are not independent: a regression that finds a relationship between two series that have nothing to do with each other, standard errors that are wrong by a known factor, and a validation split that reports a model more than five times better than it is.

Experimentation and A/B Testing

What an online experiment reports when it is too small, watched too often, or read across too many metrics: an effect inflated 2.4 times, a false positive rate of 19% instead of 5%, and a winning segment in almost half of all experiments where nothing happened.

Recommender Systems

Two fitted offsets that deliver two thirds of the accuracy gain before any latent factor is learned, a model 1.28 times worse for the users who have told it least, and the blind spot that opens when a system only ever sees ratings for what it chose to show.

Anomaly Detection

A detector that never fires scores 99.5% accuracy, a ROC of 0.9468 hides an alert queue that is 64% false, distance from the mean scores below chance when the anomalies sit at the centre, and twenty anomalies that group together hide each other from the method built to find them.

Encyclopedia (28)

Overfitting

When a model learns the noise and idiosyncrasies of its training data rather than the underlying pattern, so it performs well in training and poorly on new data.

Bias-Variance Trade-off

The decomposition of a model’s expected prediction error into bias, variance, and irreducible noise, and the tension that reducing one of the first two typically increases the other.

Cross-Validation

A resampling method that estimates a model’s test error by repeatedly fitting it on part of the data and evaluating it on the part held out.

Regularization

Any technique that constrains a model’s effective complexity in order to reduce variance and improve generalization, typically by penalizing large parameter values.

Bayes’ Theorem

A rule for updating the probability of a hypothesis in light of new evidence, by inverting a conditional probability.

Maximum Likelihood Estimation

A method of fitting a model by choosing the parameter values that make the observed data most probable.

Linear Regression

A model that predicts a numeric response as a weighted sum of the predictors, fitted by minimizing squared error.

Logistic Regression

A classification model that predicts the probability of a class by passing a linear combination of predictors through the logistic function.

Bagging and Random Forests

Ensemble methods that reduce variance by averaging many models fitted to bootstrap resamples, with random forests additionally decorrelating the trees by restricting the features available at each split.

Spline

A piecewise polynomial joined at chosen points called knots, constrained so that the function and its lower derivatives stay continuous there, giving local flexibility without the wild behaviour of a high-degree polynomial.

k-Means Clustering

An unsupervised algorithm that partitions observations into k groups by alternately assigning points to the nearest centroid and recomputing the centroids.

Expectation–Maximization

An iterative method for maximum-likelihood estimation when some variables are unobserved: it computes the posterior distribution over the hidden variables under the current parameters, then refits the parameters as though those expected counts had been observed.

k-Nearest Neighbours

A nonparametric classifier that predicts the class of a point by taking a majority vote among the k training observations closest to it.

Linear Discriminant Analysis

A generative classifier that models each class as a Gaussian and inverts those models with Bayes’ theorem; assuming one covariance matrix shared by all classes gives a linear decision boundary, and one per class gives a quadratic one.

ROC Curve

A plot of a classifier’s true positive rate against its false positive rate as the decision threshold is swept across its whole range, summarising every available trade between the two kinds of error.

Hierarchical Clustering

An unsupervised method that builds a tree of nested clusters by repeatedly fusing the two least dissimilar groups, so that cutting the tree at any height yields a clustering.

Principal Component Analysis

A technique that re-expresses data in new uncorrelated coordinates ordered by how much variance each explains, allowing dimension reduction by keeping only the first few.

p-value

The probability of observing data at least as extreme as the data in hand, computed under the assumption that the null hypothesis is true. It measures how unusual the sample would be in a world where the effect is absent, and nothing else.

Confidence Interval

A range computed from data by a procedure that, repeated over many samples, contains the true value a stated proportion of the time. The stated proportion is a property of the procedure, not of any particular interval it produces.

Statistical Power

The probability that a test rejects the null hypothesis when a specified alternative is true. It is the chance of finding an effect that is genuinely there, and it is fixed by the design before any data are collected.

Confounding

A variable that influences both the treatment and the outcome, so that a comparison between the treated and the untreated measures the difference between the groups as well as the effect of the treatment.

Causal Graph

A drawing of assumed cause-and-effect relationships as arrows between variables, used to decide which variables must be adjusted for and which must not - a question the data alone cannot answer.

Stationarity

A property of a series whose statistical behaviour does not depend on when you look at it: the mean, the variance and the correlation structure are the same in every window. Almost every classical method assumes it, and most real series lack it.

Autocorrelation

The correlation of a series with a lagged copy of itself, measuring how long the influence of an observation persists. It is the structure that makes time-series data informative and the reason ordinary standard errors do not apply to it.

Multiple Comparisons

The inflation of false positives that occurs whenever more than one test, metric, segment or stopping point is allowed to produce the headline. Each additional chance raises the probability that something crosses the threshold by luck alone.

Matrix Factorisation

A model that explains a sparse table of interactions as the product of two small matrices, giving every user and every item a short vector of learned traits whose dot product predicts the missing entries.

Precision and Recall

Two rates that split what accuracy hides: precision is the share of predicted positives that are real, and recall is the share of real positives that were found.

Anomaly Detection

Finding the few observations that were not produced by the process that produced the rest. The defining difficulty is not the algorithm but the base rate: at 0.5% anomalies, a detector that never fires is 99.5% accurate, and most standard metrics inherit that number rather than measuring skill.

Articles (21)

The Direction That Changes When You Change Units

Twelve people, two measurements, and three different first principal components: in millimetres the answer is almost pure height, in metres almost pure weight, and in centimetres an even blend - with the correlation fixed at 0.9500 throughout. What that says about what PCA maximises, why a proportion of variance explained of 99.999% can be a statement about metres rather than about people, and what standardising actually chooses.

The Control Variable That Invents a Relationship

Two independent causes and one common effect. Adjust for the effect and the causes acquire a correlation of exactly -1: a regression of A on B recovers a coefficient of +0.0030, and adding the common effect as a control turns it into -1.0000. Selecting a sample does the same thing invisibly, which is why "control for everything you measured" is not a defensible rule.

The 95% Interval That Covers 81% of the Time

The textbook confidence interval for a proportion has exact coverage you can compute by summing over the n+1 possible samples, and at n = 30 with p = 0.10 it is 0.8085 rather than 0.95. Coverage does not improve monotonically with n, and in a rare-event setting it can fall to 0.0392. Two one-line alternatives fix it.

A Score That Loses to Doing Nothing

A five-nearest-neighbour model scores 0.9983 under random five-fold cross-validation on a random walk, a series whose increments are by construction unpredictable. Evaluated forward in time it scores 0.6559 with an RMSE 12.44 times larger, and loses to carrying the last observed value forward. The split, not the model, produced the first number.

The Detector That Never Fires Is 99.5% Accurate

At a realistic base rate the do-nothing detector wins on accuracy, a ROC of 0.9468 hides an alert queue that is 64% false, distance from the mean scores below chance when anomalies sit at the centre, and twenty anomalies that group together hide each other from the method built to find them.

The Treatment That Helps Everyone and Harms the Average

A treatment that raises recovery by exactly five points in every subgroup while appearing to lower it overall, why more data makes that conclusion more confident rather than more correct, what randomisation buys that adjustment cannot, and the case where controlling for a variable manufactures an association from nothing.

The Regression That Finds a Relationship That Is Not There

Two series generated from separate random numbers come out significantly related 82.8% of the time, a standard error on dependent data is too narrow by a computable factor of 2.4, and the usual validation split reports a forecaster more than five times better than it is. Three failures, one cause, and the checks that catch each of them.

The Model That Picks Its Own Training Data

Two fitted offsets deliver 66% of a recommender’s gain in accuracy before any latent factor is learned, the error is 1.28 times worse for the users who have said least, only 30% of the catalogue reaches anyone’s top ten with no explicit popularity term in the model, and after six rounds of self-selected data the system is 1.14 times worse exactly where it stopped looking.

The Experiment That Was Going to Win Anyway

A test with 2,000 users per arm reports effects 2.4 times too large. An A/A test checked ten times comes out significant 19% of the time. Twenty independent null metrics produce a winner 64% of the time and twelve null segments 46%. Four numbers, one cause, and the decisions that have to be made before the data arrives.

What a Sample Can and Cannot Tell You

Estimators as random variables with distributions of their own, the case where the unbiased estimator is the worse one, what a confidence interval actually promises and the standard interval that delivers 87% where it advertises 95%, and what a p-value is a probability of - every figure computed exactly or by fixed-seed simulation.

Learning the Numbers in a Probability Model

Where the numbers in a Bayesian network or a Gaussian actually come from: the three-step maximum-likelihood recipe worked through on discrete and continuous parameters, the Beta prior that repairs what it does to an unseen event, naive Bayes and the single zero count that destroys it, and the EM algorithm for the case where the counts cannot be taken at all - with every figure computed rather than asserted.

Comparing Classifiers, and What Accuracy Hides

The Bayes classifier nothing can beat and the error floor it leaves behind, k-nearest-neighbours as a nonparametric imitation with k as the flexibility dial, discriminant analysis and why a shared covariance forces a straight line, and the confusion matrix, thresholds and ROC curve that a single accuracy figure conceals - every number computed on simulated data where the optimum is known.

Moving Beyond Linearity: Splines and Additive Models

How to fit curved relationships without leaving least squares: basis functions, the constraints that turn a broken piecewise polynomial into a spline, the single extra column per knot that enforces them for free, and the roughness penalty that lets a curve choose its own flexibility.

Unsupervised Learning: Structure Without Labels

What changes when there is no response to predict: principal components as the direction of maximum variance, K-means and the local optima it settles into, hierarchical clustering and the linkage that decides the answer - and why none of the required choices can be validated the way a classifier can.

What Is Statistical Learning?

The setup behind every predictive model: estimating an unknown function f from data, the split between reducible and irreducible error, and why prediction and inference pull in different directions.

The Bias-Variance Tradeoff

The exact decomposition of expected test error into squared bias, variance, and irreducible noise, demonstrated numerically with a 2,000-run simulation where all three terms are measured separately and shown to add up.

Cross-Validation and Resampling

Why training error is a biased estimate of test error, and how the validation set, leave-one-out, and k-fold approaches fix it, with a five-fold LOOCV computation worked out observation by observation.

Linear Regression from First Principles

Derive the least-squares coefficients by differentiating the residual sum of squares, then work a complete five-observation fit by hand: coefficients, fitted values, residuals, RSS, and R-squared, each verified numerically.

Logistic Regression and Classification

Why a straight line cannot model a probability, how the logistic function fixes it, and what the coefficients mean in log-odds, with a gradient-ascent step and a converged fit computed and checked numerically.

Regularization: Ridge and Lasso

Adding a penalty on coefficient size to trade a little bias for a large reduction in variance, and why the L1 penalty sets coefficients exactly to zero while L2 only shrinks them, with both fitted numerically.

Decision Trees and Ensembles

How recursive binary splitting builds a tree, why the Gini index beats accuracy as a splitting criterion, and how bagging and random forests turn a high-variance learner into a strong one, with the split arithmetic worked out.

Tools (7)

Datasets (3)

Research (2)

Projects (2)

Related topics