Skip to content
Kudos AI

What Accuracy Hides

At a 0.5% base rate the do-nothing detector wins on accuracy, ROC reports 0.9468 for a model whose top hundred alerts are 64% false, and precision@k is the only number that describes the queue somebody has to work through.

IntermediateModule 125 min · 100 XP
An accuracy figure of 99.5% sitting above a detector that never fires, then the same detector scored twice - a ROC curve hugging the corner while the precision-recall curve collapses.

Twenty thousand events. One hundred of them are anomalies. Build a detector that returns "normal" for everything.

It is 99.50% accurate.

That is the whole problem in one line, and every metric habit carried over from balanced classification has to be re-examined against it.

Two scores for one detector

Now take a detector that genuinely works: it ranks anomalies higher than normal events, with the overlap any real scorer has. Score it two ways.

metricvalue
ROC AUC0.9468
average precision (PR AUC)0.3245
what a random detector scores on the second0.0050

Both numbers are correctly computed from the same ranking. They disagree because they divide by different things.

The false-positive rate divides by the 19,900 normal events. A thousand false alarms move it by five percentage points, so the ROC curve stays near the corner and the AUC reads like an excellent model.

Precision divides by the number of alerts. A thousand false alarms are not a rounding error there; they are the entire queue.

So the question a metric answers matters more than its value. ROC answers "how well does this rank anomalies above normals", which is a property of the detector. Precision answers "what fraction of what I am shown is real", which is what the person on call experiences.

The number that describes the queue

Neither AUC tells anyone how many things to look at. Precision@k does, because it fixes the queue length first:

you reviewof those, realprecisionrecall
top 502652%26%
top 1003636%36%
top 5006713.4%67%

At the top hundred, somebody reviews 100 items to find 36. Nearly two false alarms for every true find - and that is the good case, from a detector whose ROC says 0.9468.

This is the number to negotiate over, because it converts directly into staffing. An analyst who can check forty items a day is choosing a row of that table, and the row determines what share of anomalies the organisation catches. A threshold chosen without reference to that capacity produces either an ignored queue or a false sense of coverage.

The figure below computes both scores from the scoring model rather than from one draw, which sharpens the exercise above. ROC does not barely move when anomalies get ten times commoner; it does not move at all, because it is a function of the two score distributions and nothing else. Average precision doubles and more. The second panel turns a review capacity into the two numbers that decide what the organisation actually sees.

Interactive: one ranking, two metrics that disagree

Both computed exactly from the scoring model, not from one draw.

ROC0
ROC AUC
0.9552
Average precision
0.3326
Random detector scores
0.0050
Never-fire accuracy
99.5%

At a contamination rate of 0.5%, a detector that returns “normal” for everything is 99.5% accurate, which is the whole problem in one number. The working detector scores 0.9552 on ROC and 0.3326 on average precision, from the same ranking. Drag the rate: ROC does not move at all, because it measures how well anomalies are ranked above normal events and that has nothing to do with how many there are. Average precision moves a lot, because precision divides by the number of alerts. A thousand false alarms are five points of false-positive rate and the entire queue.

How to report a detector

Three things, and the third is the one usually missing:

  1. the base rate of the data it was measured on, because every other number depends on it
  2. precision and recall at a stated k, not an AUC alone
  3. the review capacity that k assumes, because a detector is only as good as the queue somebody actually works

"ROC 0.95" is not a claim anyone can act on. "At forty reviews a day we catch a quarter of incidents, and three in eight of what we show you is noise" is.

What this sets up

This lesson assumed a detector and measured it. The next asks where the score comes from - and finds that the three main families disagree about what "anomalous" even means, in ways that show up as one of them scoring below chance.

References & further reading

  • Charu C. Aggarwal, Recommender Systems: The Textbook, Springer, 2016· Kudos AI reference library
  • Kevin P. Murphy, Probabilistic Machine Learning: An Introduction, MIT Press (Adaptive Computation and Machine Learning), 2022source ↗

Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.

Unlock the full path

This first lesson is free. Enrol to take the mastery quiz, earn XP, and unlock every module, with more interactive, runnable examples throughout.