Skip to content
Kudos AI
Lire en français
Anomaly Detection

The Detector That Never Fires Is 99.5% Accurate

At a realistic base rate the do-nothing detector wins on accuracy, a ROC of 0.9468 hides an alert queue that is 64% false, distance from the mean scores below chance when anomalies sit at the centre, and twenty anomalies that group together hide each other from the method built to find them.

7 min readKudos AI

Prerequisites: Classification Methods Compared

An accuracy figure of 99.5% above a detector that never fires, a ROC curve hugging the corner while the precision-recall curve collapses, and a growing clump of anomalies making each other look ordinary.

Twenty thousand events, a hundred of them anomalies. Return "normal" for everything and you are 99.50% accurate.

That is not a joke about metrics. It is the constraint the whole subject is built around, and it survives every attempt to route past it with a better model.

Every metric inherits the base rate

Take a detector that genuinely works - it ranks anomalies above normal events, with the overlap any real scorer has. Score it twice:

metricvalue
ROC AUC0.9468
average precision0.3245
what a random detector scores on the second0.0050

Both are correct. They divide by different things.

The false-positive rate divides by the 19,900 normal events, so a thousand false alarms shift it by five points and the ROC curve stays glued to the corner. Precision divides by the number of alerts, where a thousand false alarms are the entire queue.

So the question a metric answers matters more than the number. ROC describes the detector. Precision describes what the person on call sees:

you reviewof those, realprecisionrecall
top 502652%26%
top 1003636%36%
top 5006713.4%67%

Nearly two false alarms per true find, from a model whose AUC says 0.9468.

Classifier metrics explorer

Move the threshold and watch precision and recall trade off.

Precision
9.8%
Recall
84.1%
F1
17.5%
Accuracy
84.1%
Majority-class baseline
98.0%

Confusion matrix

Predicted +Predicted −
Actual +16832
Actual −1,5558,245

ROC AUC: 0.921

ROC curve

FPRTPR

Raising the threshold buys precision by giving up recall, and lowering it does the reverse. On a rare positive class, accuracy is nearly useless: it is dominated by the negatives, so a model that never fires still scores well. Precision and recall are the honest summary.

The figure below computes both scores from the scoring model rather than from one draw, which sharpens the exercise above. ROC does not barely move when anomalies get ten times commoner; it does not move at all, because it is a function of the two score distributions and nothing else. Average precision doubles and more. The second panel turns a review capacity into the two numbers that decide what the organisation actually sees.

Interactive: one ranking, two metrics that disagree

Both computed exactly from the scoring model, not from one draw.

ROC0
ROC AUC
0.9552
Average precision
0.3326
Random detector scores
0.0050
Never-fire accuracy
99.5%

At a contamination rate of 0.5%, a detector that returns “normal” for everything is 99.5% accurate, which is the whole problem in one number. The working detector scores 0.9552 on ROC and 0.3326 on average precision, from the same ranking. Drag the rate: ROC does not move at all, because it measures how well anomalies are ranked above normal events and that has nothing to do with how many there are. Average precision moves a lot, because precision divides by the number of alerts. A thousand false alarms are five points of false-positive rate and the entire queue.

Three families, three blind spots

Where does the score come from? Build data with two kinds of anomaly - ten points far above everything, and twenty in the sparse gap between two dense clusters, where nothing normal lives but which sits in the middle of the data. Then run four standard detectors. The two right-hand columns are recall inside the top 60 of each ranking, twice the 30 anomalies planted:

detectorROC AUCof the 10 far pointsof the 20 gap points
Mahalanobis distance0.335100%0%
kNN distance0.978100%65%
local outlier factor0.933100%30%
isolation forest0.903100%0%

Mahalanobis scores below chance. Nothing is broken: the gap sits at the global mean, so those twenty points have the lowest distance scores in the dataset and rank below ordinary data. The method encodes a definition - an anomaly is far from the middle - and this data violates it.

Isolation forest manages a respectable 0.903 and reaches none of the gap inside that top 60, because a point at the centre of both coordinate ranges takes nearly as many axis-parallel cuts to isolate as any ordinary point. Nearly, not exactly: every gap point still outranks at least 73% of the normal data, gap against normal at ROC 0.855, and pushing those twenty to the bottom instead would drop the overall figure to 0.333. The margin is thin, which is a different failure from having none.

The local method is defeated by a crowd

The local outlier factor exists for exactly this case: it compares a point's density to its neighbours', so a point in a sparse pocket stands out even at the centre. It found 30%. Grow the group and watch why:

anomalies in the gaplocal outlier factorkNN distance
2100%100%
560%80%
1060%100%
2040%80%
408%25%

This is masking. LOF's premise is that a point's neighbours are normal. Once forty anomalies sit together they are each other's neighbourhood, and the comparison reports a perfectly ordinary local density.

Plain kNN distance degrades too, but far more slowly, because absolute distance does not care whether the neighbours are anomalous. It is the unusual case where the simpler method is the more robust one - worth remembering when a sophisticated detector is proposed for a problem where anomalies arrive in bursts. Fraud, outages and sensor faults all arrive in bursts.

The figure below is an independent implementation of that dataset with an independent generator, and Mahalanobis still reads 0.336 against the 0.335 above: scoring below chance here is a property of the geometry, not of anyone's seed. Switch detectors to see which points each one flags, then drag the bridge from two anomalies to forty. The local outlier factor falls from 0.999 to 0.62 while kNN distance only slips to 0.95, which is masking measured rather than described.

Interactive: three definitions of anomalous, one dataset

The flagged points are the top of each ranking, on a budget of one per anomaly.

ROC AUC
0.336
Of the far points
100%
Of the bridge
0%
Points scored
830

Mahalanobis reads 0.336 - worse than a coin, and nothing is broken. It measures distance from the centre, the bridge sits at the centre, so those 20 points have the LOWEST scores in the dataset: ranked below ordinary data rather than above it. It still finds 100% of the far points and 0% of the bridge. The method encodes a definition - an anomaly is far from the middle - and this data contains anomalies that are not. A detector cannot find a kind of anomaly its score function does not express.

The threshold controls recall; contamination controls precision

Price the two mistakes - say a miss costs 500 and a false alarm 20 - and the threshold stops being a matter of taste:

policyalertscaughttotal cost
never fire00 / 10050,000
flag the top 1%20152 / 10026,980
cost-minimising46067 / 10024,360

The top-1% rule lands within 11% of optimal here, which is luck rather than principle: it fires less than half the alerts it should, and at a different cost ratio it would be badly wrong.

Then leave that threshold alone and let the world move:

anomaly ratealertsprecisionrecall
half the usual4258%64%
what it was tuned on46015%67%
twice the usual52325%65%

Recall barely moves; precision swings threefold. Recall is the share of anomalies above the cut, so it depends on the anomaly score distribution alone, and that distribution has not changed - only the number of draws from it. The normal distribution has not changed either, which holds the false alarms near 2.3% of the normal events. So precision is a moving count of real anomalies over a nearly fixed pile of false alarms, and that ratio is what swings.

A team promised "most alerts will be real" will see that promise broken by a quiet month, with no change to the model, the pipeline or the threshold.

The figure below lets you do the thing the table only describes: hold the cut still and move the world. Drag the contamination and the recall read-out does not change by a single point, because recall reads off the anomaly score distribution alone and that distribution has not moved. Precision changes about thirteen-fold across the slider's full range, from 10 to 200 anomalies per 10,000. The capacity table underneath is the same detector reported the way it should be

  • given how many items can be reviewed in a day, here is what that buys.

Interactive: hold the cut still and let the world move

The threshold never changes what share of anomalies it catches.

2.0
normal eventsanomalies
Alerts
518
Precision
13%
Recall
66%
Accuracy
97.56%

What each review capacity actually buys

reviewed per dayprecisionrecall
2073%15%
10037%37%
46014%63%
10008%76%

20,000 events

At this cut the detector catches 66% of whatever anomalies exist, and that number does not depend on how many there are - drag the contamination slider and watch it refuse to move. Precision is 13% and it moves with the rate almost proportionally, because the alerts are overwhelmingly false ones and it is their ratio to the real ones that changes. So a team promised that most of what they see will be real has that promise broken by a quiet month, with no change to the model, the data or the threshold. Note the accuracy read-out while you do it: at 50.0 anomalies per ten thousand events, a detector that never fired at all would score about the same.

Specify the problem before choosing a method

Three numbers decide whether an anomaly detector is worth anything, and all three are known before any model is trained:

  1. the base rate - it sets what every metric can mean
  2. the review capacity - it fixes k, and k fixes the threshold, not the other way round
  3. the cost of a miss relative to a false alarm - it turns the threshold into arithmetic

Report a detector as "at forty reviews a day we catch a quarter of incidents, and three in eight of what we show you is noise". Not as "ROC 0.95".

And check which kind of anomaly your data actually contains, because every family is blind to a kind another one finds easily. Running two families with different blind spots and looking at what only one of them flags is cheap - and the disagreement is usually where the interesting anomaly is.

References & further reading

  • Charu C. Aggarwal, Recommender Systems: The Textbook, Springer, 2016· Kudos AI reference library
  • Kevin P. Murphy, Probabilistic Machine Learning: An Introduction, MIT Press (Adaptive Computation and Machine Learning), 2022source ↗

Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.

Related reading

6 min readUnsupervised Learning

The Direction That Changes When You Change Units

Twelve people, two measurements, and three different first principal components: in millimetres the answer is almost pure height, in metres almost pure weight, and in centimetres an even blend - with the correlation fixed at 0.9500 throughout. What that says about what PCA maximises, why a proportion of variance explained of 99.999% can be a statement about metres rather than about people, and what standardising actually chooses.

Machine LearningStatistics
4 min readTime Series

A Score That Loses to Doing Nothing

A five-nearest-neighbour model scores 0.9983 under random five-fold cross-validation on a random walk, a series whose increments are by construction unpredictable. Evaluated forward in time it scores 0.6559 with an RMSE 12.44 times larger, and loses to carrying the last observed value forward. The split, not the model, produced the first number.

StatisticsMachine Learning
8 min readCausal Inference

The Treatment That Helps Everyone and Harms the Average

A treatment that raises recovery by exactly five points in every subgroup while appearing to lower it overall, why more data makes that conclusion more confident rather than more correct, what randomisation buys that adjustment cannot, and the case where controlling for a variable manufactures an association from nothing.

StatisticsMachine Learning
← Back to all articles