What Is Statistical Learning?
The setup behind every predictive model: estimating an unknown function f from data, the split between reducible and irreducible error, and why prediction and inference pull in different directions.
المربّعات الصغرى والانحدار اللوجستي وريدج ولاسو والتحقق المتقاطع k-fold، مُنفَّذة من معادلات التقدير ومُدقَّقة مقابل scikit-learn.
مكتبة مُوجزة تُنفّذ صميم التعلّم الإحصائي مباشرةً من الرياضيات: المعادلات الناظمية للمربّعات الصغرى، والمربّعات الصغرى المُعاد ترجيحها تكراريًا للانحدار اللوجستي، والحل المغلق لريدج، والنزول الإحداثي للاسو. يُتحقَّق من كل مُقدِّر عدديًا مقابل نظيره في scikit-learn، فيرى القارئ أن اشتقاقًا أُجري على الورق يُنتج المعاملات نفسها التي يُنتجها التنفيذ الإنتاجي. وفوق المُقدِّرات تقع طبقة إعادة المعاينة - التحقق المتقاطع k-fold وleave-one-out - تُستخدم لرسم منحنى التحيّز والتباين تجريبيًا: فكلما ارتفعت مرونة النموذج انخفض خطأ التدريب رتيبًا بينما ينعطف خطأ التحقق صعودًا، وترسم المكتبة بالضبط موضع هذا الانعطاف. التنفيذ جارٍ ولم يُنشر بعد أي مستودع للمصدر.
The setup behind every predictive model: estimating an unknown function f from data, the split between reducible and irreducible error, and why prediction and inference pull in different directions.
Derive the least-squares coefficients by differentiating the residual sum of squares, then work a complete five-observation fit by hand: coefficients, fitted values, residuals, RSS, and R-squared, each verified numerically.
Why a straight line cannot model a probability, how the logistic function fixes it, and what the coefficients mean in log-odds, with a gradient-ascent step and a converged fit computed and checked numerically.
Adding a penalty on coefficient size to trade a little bias for a large reduction in variance, and why the L1 penalty sets coefficients exactly to zero while L2 only shrinks them, with both fitted numerically.
Why training error is a biased estimate of test error, and how the validation set, leave-one-out, and k-fold approaches fix it, with a five-fold LOOCV computation worked out observation by observation.
The exact decomposition of expected test error into squared bias, variance, and irreducible noise, demonstrated numerically with a 2,000-run simulation where all three terms are measured separately and shown to add up.