What Is Statistical Learning?
The setup behind every predictive model: estimating an unknown function f from data, the split between reducible and irreducible error, and why prediction and inference pull in different directions.
Draw thirty samples, form a confidence interval from each, and count how many cover the truth - the clearest way to see that the confidence level describes the procedure, not any single interval.
30 samples, each a 95% CI.
Each red interval missed the true value. Over many samples, close to 95% of the intervals cover it, that is what the confidence level means: a property of the procedure, not a probability that this one interval contains the truth (it either does or does not). Resample to watch which intervals miss change, and shrink the sample size to see the intervals widen.
Runs entirely in your browser. Nothing you enter is uploaded or stored.
The setup behind every predictive model: estimating an unknown function f from data, the split between reducible and irreducible error, and why prediction and inference pull in different directions.
Why training error is a biased estimate of test error, and how the validation set, leave-one-out, and k-fold approaches fix it, with a five-fold LOOCV computation worked out observation by observation.