What Is Statistical Learning?
The setup behind every predictive model: estimating an unknown function f from data, the split between reducible and irreducible error, and why prediction and inference pull in different directions.
Least squares, logistic regression, ridge and lasso, and k-fold cross-validation implemented from their estimating equations and checked against scikit-learn.
A compact library implementing the core of statistical learning directly from the mathematics: the normal equations for least squares, iteratively reweighted least squares for logistic regression, the closed-form ridge solution, and coordinate descent for the lasso. Each estimator is validated numerically against its scikit-learn counterpart, so a reader can see that a derivation carried out on paper produces the same coefficients as the production implementation. On top of the estimators sits a resampling layer - k-fold and leave-one-out cross-validation - used to draw the bias-variance curve empirically: as model flexibility rises, training error falls monotonically while validation error turns upward, and the toolkit plots exactly where that turn happens. Implementation is in progress and no source repository has been published yet.
The setup behind every predictive model: estimating an unknown function f from data, the split between reducible and irreducible error, and why prediction and inference pull in different directions.
Derive the least-squares coefficients by differentiating the residual sum of squares, then work a complete five-observation fit by hand: coefficients, fitted values, residuals, RSS, and R-squared, each verified numerically.
Why a straight line cannot model a probability, how the logistic function fixes it, and what the coefficients mean in log-odds, with a gradient-ascent step and a converged fit computed and checked numerically.
Adding a penalty on coefficient size to trade a little bias for a large reduction in variance, and why the L1 penalty sets coefficients exactly to zero while L2 only shrinks them, with both fitted numerically.
Why training error is a biased estimate of test error, and how the validation set, leave-one-out, and k-fold approaches fix it, with a five-fold LOOCV computation worked out observation by observation.
The exact decomposition of expected test error into squared bias, variance, and irreducible noise, demonstrated numerically with a 2,000-run simulation where all three terms are measured separately and shown to add up.