Skip to content
Kudos AI

Precision and Recall

Two rates that split what accuracy hides: precision is the share of predicted positives that are real, and recall is the share of real positives that were found.

Also known as: Positive predictive value and sensitivity, F1 score

Understanding Precision and Recall

A confusion matrix has four cells, and accuracy collapses them into one number by counting the diagonal. When one class is rare that number is dominated by the true negatives, and a model that never predicts the rare class scores well on it. Precision and recall cut the matrix the other way: precision divides the true positives by everything predicted positive, recall divides them by everything actually positive. Neither uses the true negatives at all.

The two fail in different directions and it is worth naming them separately. Low recall means missing real cases, which is what matters in screening, search and fraud detection. Low precision means crying wolf, which is what matters when each alarm costs a person’s attention. A system with both problems is easy to spot; a system with one is easy to mistake for a good one if only the other is reported.

Neither is a fixed property of a model that outputs a score. Lowering the threshold catches more real positives and raises recall while admitting more false alarms and lowering precision. The curve traced out as the threshold moves is the honest description; a single pair of numbers is a point on it, chosen deliberately or by default.

The F1 score summarises the pair as a harmonic mean rather than an arithmetic one, which refuses to be rescued by a single high value: a precision of 1 with a recall of 0.02 has an arithmetic mean of 0.51 and an F1 of 0.039. Where the two errors have genuinely different costs, weighting them explicitly says more than any single summary.

How to Calculate

\text{precision} = \frac{TP}{TP + FP}, \qquad \text{recall} = \frac{TP}{TP + FN}, \qquad F_1 = \frac{2\,\text{precision}\cdot\text{recall}}{\text{precision} + \text{recall}}

where

TP
true positives: real cases correctly flagged
FP
false positives: false alarms
FN
false negatives: real cases missed

Example of Precision and Recall

Screen 20,000 people for a condition that 1 in 200 has, so 100 are genuinely positive and 19,900 are not. A detector that never fires is right 19,900 times out of 20,000: an accuracy of 0.995, and a recall of 0.

Now take a detector with 90 per cent sensitivity and 95 per cent specificity. It finds 90 of the 100 real cases and raises 995 false alarms, giving a recall of 0.9 and a precision of 90/1085 = 0.082949. Eleven of every twelve alarms are false, and the false alarms outnumber the true ones 11.06 to one.

Its accuracy is 18,995/20,000 = 0.94975, which is LOWER than the detector that never fires. Accuracy prefers the useless model; precision and recall say plainly what each one does. This is why the pair is reported for rare classes and accuracy is not.

Frequently Asked Questions

Which one should be optimised?

Whichever error is more expensive, which is a question about the application and not about the data. Missing a tumour and flagging a healthy patient are not interchangeable, and no summary statistic knows the exchange rate.

Why does precision fall so far when the class is rare?

Because false alarms are drawn from the large negative pool. At 1 in 200 prevalence even a 5 per cent false positive rate produces ten times as many alarms as there are real cases, whatever the sensitivity is.

The Bottom Line

Precision asks how many alarms were real, recall asks how many real cases were caught, and neither counts the true negatives. On a rare class that is the difference between a number that flatters a useless model and a pair that describes what it actually does.