Understanding ROC Curve
A classifier that outputs a score becomes a decision rule only once a threshold is fixed. The default of one half is a convention, not a result, and it embeds the assumption that a false positive and a false negative are equally costly. Wherever that is untrue - and it usually is - the threshold is a parameter of the problem rather than of the model, and it should be chosen deliberately.
The confusion matrix is the first step away from a single number: it separates the errors into false positives and false negatives, from which sensitivity (the proportion of true positives correctly identified) and specificity (the same for true negatives) follow. Moving the threshold moves both, in opposite directions, and the ROC curve is simply the record of every such pair as the threshold sweeps from one extreme to the other.
A perfect classifier reaches the top-left corner, where every positive is caught and no negative is raised; a classifier with no information traces the diagonal. The area under the curve condenses the whole picture into one number that has an exact probabilistic reading: it is the probability that a randomly chosen positive receives a higher score than a randomly chosen negative. The equivalence to the Mann–Whitney U statistic is an identity, not an approximation.
What the measure deliberately ignores is worth being explicit about. Because only the ordering of the scores matters, AUC is unchanged by any monotone transformation of them, which makes it useful for comparing models before a threshold has been chosen and useless as a check on whether the predicted probabilities can be believed. On heavily imbalanced problems the precision-recall curve often shows differences that the ROC curve flattens.
How to Calculate
AUC = P(score(X⁺) > score(X⁻)), sensitivity = TP/(TP+FN), specificity = TN/(TN+FP)
where
- X⁺, X⁻
- a randomly chosen positive and a randomly chosen negative observation
- TP, FN
- true positives and false negatives - positives caught and positives missed
- TN, FP
- true negatives and false positives - negatives cleared and false alarms
- threshold
- the score above which a prediction is called positive; sweeping it traces the curve
Example of ROC Curve
A quadratic discriminant classifier on 200,000 test points has an overall error rate of 0.093480 at the default threshold, which conceals a genuine asymmetry: sensitivity 0.961083 against specificity 0.851753. Checking the optimal rule for the same problem gives 0.952845 and 0.861738, so the imbalance is a property of the problem rather than a defect of the fit.
Lowering the threshold from 0.5 to 0.1 raises sensitivity to 0.995928 and drops specificity to 0.750784, while the overall error rate rises from 0.093480 to 0.126415. The model is identical throughout - only the threshold moved - and whether the trade is worthwhile depends entirely on the relative cost of a missed positive and a false alarm.
Summarised over all thresholds, this classifier has an AUC of 0.965269 against 0.933488 for a linear discriminant classifier on the same data. Computing that area by the trapezoid rule on the ROC curve and computing the Mann–Whitney statistic on the raw scores agree exactly.
Frequently Asked Questions
When is accuracy the wrong summary?
Whenever the two kinds of error carry different costs, and whenever the classes are strongly imbalanced - a classifier that always predicts the majority class can achieve high accuracy while being useless. The confusion matrix should be the default report, with accuracy derived from it rather than quoted alone.
How should the threshold be chosen?
From the costs, not from the data. If a missed positive is ten times as expensive as a false alarm, the threshold that minimises expected cost is far from one half. Where costs cannot be quantified, an explicit target for sensitivity or specificity is usually a more honest constraint than minimising the error rate.
Is a higher AUC always the better model?
Not necessarily. AUC averages over thresholds that may all be irrelevant to the intended use, so a model with lower AUC can be preferable if it performs better in the region of the curve that will actually be operated in. And AUC says nothing about whether the predicted probabilities themselves are trustworthy.
The Bottom Line
Report the confusion matrix before the accuracy, treat the threshold as a decision about costs rather than a default, and read the ROC curve as the menu of trades available from one fixed model. AUC ranks; it does not calibrate.