Logistic Regression and Classification
Why a straight line cannot model a probability, how the logistic function fixes it, and what the coefficients mean in log-odds, with a gradient-ascent step and a converged fit computed and checked numerically.
Move a decision threshold across an imbalanced population and watch precision, recall, F1, and the ROC point move with it - including the regime where accuracy looks excellent and the model is useless.
Move the threshold and watch precision and recall trade off.
| Predicted + | Predicted − | |
|---|---|---|
| Actual + | 168 | 32 |
| Actual − | 1,555 | 8,245 |
ROC AUC: 0.921
Raising the threshold buys precision by giving up recall, and lowering it does the reverse. On a rare positive class, accuracy is nearly useless: it is dominated by the negatives, so a model that never fires still scores well. Precision and recall are the honest summary.
Runs entirely in your browser. Nothing you enter is uploaded or stored.
Why a straight line cannot model a probability, how the logistic function fixes it, and what the coefficients mean in log-odds, with a gradient-ascent step and a converged fit computed and checked numerically.
The setup behind every predictive model: estimating an unknown function f from data, the split between reducible and irreducible error, and why prediction and inference pull in different directions.