What Is Statistical Learning?
The setup behind every predictive model: estimating an unknown function f from data, the split between reducible and irreducible error, and why prediction and inference pull in different directions.
Moindres carrés, régression logistique, ridge et lasso, et validation croisée k-fold, implémentés depuis leurs équations d’estimation et vérifiés face à scikit-learn.
Une bibliothèque compacte implémentant le cœur de l’apprentissage statistique directement à partir des mathématiques : équations normales pour les moindres carrés, moindres carrés repondérés itérativement pour la régression logistique, solution en forme close pour la ridge, et descente par coordonnées pour le lasso. Chaque estimateur est validé numériquement face à son équivalent scikit-learn, de sorte qu’un lecteur constate qu’une dérivation menée sur papier produit les mêmes coefficients que l’implémentation de production. Au-dessus des estimateurs se trouve une couche de rééchantillonnage - validation croisée k-fold et leave-one-out - utilisée pour tracer empiriquement la courbe biais-variance : à mesure que la flexibilité du modèle augmente, l’erreur d’entraînement décroît de façon monotone tandis que l’erreur de validation finit par remonter, et la bibliothèque trace exactement où ce retournement se produit. L’implémentation est en cours et aucun dépôt source n’a encore été publié.
The setup behind every predictive model: estimating an unknown function f from data, the split between reducible and irreducible error, and why prediction and inference pull in different directions.
Derive the least-squares coefficients by differentiating the residual sum of squares, then work a complete five-observation fit by hand: coefficients, fitted values, residuals, RSS, and R-squared, each verified numerically.
Why a straight line cannot model a probability, how the logistic function fixes it, and what the coefficients mean in log-odds, with a gradient-ascent step and a converged fit computed and checked numerically.
Adding a penalty on coefficient size to trade a little bias for a large reduction in variance, and why the L1 penalty sets coefficients exactly to zero while L2 only shrinks them, with both fitted numerically.
Why training error is a biased estimate of test error, and how the validation set, leave-one-out, and k-fold approaches fix it, with a five-fold LOOCV computation worked out observation by observation.
The exact decomposition of expected test error into squared bias, variance, and irreducible noise, demonstrated numerically with a 2,000-run simulation where all three terms are measured separately and shown to add up.