The Detector That Never Fires Is 99.5% Accurate
At a realistic base rate the do-nothing detector wins on accuracy, a ROC of 0.9468 hides an alert queue that is 64% false, distance from the mean scores below chance when anomalies sit at the centre, and twenty anomalies that group together hide each other from the method built to find them.
Prerequisites: Classification Methods Compared
Twenty thousand events, a hundred of them anomalies. Return "normal" for everything and you are 99.50% accurate.
That is not a joke about metrics. It is the constraint the whole subject is built around, and it survives every attempt to route past it with a better model.
Every metric inherits the base rate
Take a detector that genuinely works - it ranks anomalies above normal events, with the overlap any real scorer has. Score it twice:
| metric | value |
|---|---|
| ROC AUC | 0.9468 |
| average precision | 0.3245 |
| what a random detector scores on the second | 0.0050 |
Both are correct. They divide by different things.
The false-positive rate divides by the 19,900 normal events, so a thousand false alarms shift it by five points and the ROC curve stays glued to the corner. Precision divides by the number of alerts, where a thousand false alarms are the entire queue.
So the question a metric answers matters more than the number. ROC describes the detector. Precision describes what the person on call sees:
| you review | of those, real | precision | recall |
|---|---|---|---|
| top 50 | 26 | 52% | 26% |
| top 100 | 36 | 36% | 36% |
| top 500 | 67 | 13.4% | 67% |
Nearly two false alarms per true find, from a model whose AUC says 0.9468.
Classifier metrics explorer
Move the threshold and watch precision and recall trade off.
- Precision
- 9.8%
- Recall
- 84.1%
- F1
- 17.5%
- Accuracy
- 84.1%
- Majority-class baseline
- 98.0%
Confusion matrix
| Predicted + | Predicted − | |
|---|---|---|
| Actual + | 168 | 32 |
| Actual − | 1,555 | 8,245 |
ROC AUC: 0.921
ROC curve
Raising the threshold buys precision by giving up recall, and lowering it does the reverse. On a rare positive class, accuracy is nearly useless: it is dominated by the negatives, so a model that never fires still scores well. Precision and recall are the honest summary.
The figure below computes both scores from the scoring model rather than from one draw, which sharpens the exercise above. ROC does not barely move when anomalies get ten times commoner; it does not move at all, because it is a function of the two score distributions and nothing else. Average precision doubles and more. The second panel turns a review capacity into the two numbers that decide what the organisation actually sees.
Interactive: one ranking, two metrics that disagree
Both computed exactly from the scoring model, not from one draw.
- ROC AUC
- 0.9552
- Average precision
- 0.3326
- Random detector scores
- 0.0050
- Never-fire accuracy
- 99.5%
At a contamination rate of 0.5%, a detector that returns “normal” for everything is 99.5% accurate, which is the whole problem in one number. The working detector scores 0.9552 on ROC and 0.3326 on average precision, from the same ranking. Drag the rate: ROC does not move at all, because it measures how well anomalies are ranked above normal events and that has nothing to do with how many there are. Average precision moves a lot, because precision divides by the number of alerts. A thousand false alarms are five points of false-positive rate and the entire queue.
Three families, three blind spots
Where does the score come from? Build data with two kinds of anomaly - ten points far above everything, and twenty in the sparse gap between two dense clusters, where nothing normal lives but which sits in the middle of the data. Then run four standard detectors. The two right-hand columns are recall inside the top 60 of each ranking, twice the 30 anomalies planted:
| detector | ROC AUC | of the 10 far points | of the 20 gap points |
|---|---|---|---|
| Mahalanobis distance | 0.335 | 100% | 0% |
| kNN distance | 0.978 | 100% | 65% |
| local outlier factor | 0.933 | 100% | 30% |
| isolation forest | 0.903 | 100% | 0% |
Mahalanobis scores below chance. Nothing is broken: the gap sits at the global mean, so those twenty points have the lowest distance scores in the dataset and rank below ordinary data. The method encodes a definition - an anomaly is far from the middle - and this data violates it.
Isolation forest manages a respectable 0.903 and reaches none of the gap inside that top 60, because a point at the centre of both coordinate ranges takes nearly as many axis-parallel cuts to isolate as any ordinary point. Nearly, not exactly: every gap point still outranks at least 73% of the normal data, gap against normal at ROC 0.855, and pushing those twenty to the bottom instead would drop the overall figure to 0.333. The margin is thin, which is a different failure from having none.
The local method is defeated by a crowd
The local outlier factor exists for exactly this case: it compares a point's density to its neighbours', so a point in a sparse pocket stands out even at the centre. It found 30%. Grow the group and watch why:
| anomalies in the gap | local outlier factor | kNN distance |
|---|---|---|
| 2 | 100% | 100% |
| 5 | 60% | 80% |
| 10 | 60% | 100% |
| 20 | 40% | 80% |
| 40 | 8% | 25% |
This is masking. LOF's premise is that a point's neighbours are normal. Once forty anomalies sit together they are each other's neighbourhood, and the comparison reports a perfectly ordinary local density.
Plain kNN distance degrades too, but far more slowly, because absolute distance does not care whether the neighbours are anomalous. It is the unusual case where the simpler method is the more robust one - worth remembering when a sophisticated detector is proposed for a problem where anomalies arrive in bursts. Fraud, outages and sensor faults all arrive in bursts.
The figure below is an independent implementation of that dataset with an independent generator, and Mahalanobis still reads 0.336 against the 0.335 above: scoring below chance here is a property of the geometry, not of anyone's seed. Switch detectors to see which points each one flags, then drag the bridge from two anomalies to forty. The local outlier factor falls from 0.999 to 0.62 while kNN distance only slips to 0.95, which is masking measured rather than described.
Interactive: three definitions of anomalous, one dataset
The flagged points are the top of each ranking, on a budget of one per anomaly.
- ROC AUC
- 0.336
- Of the far points
- 100%
- Of the bridge
- 0%
- Points scored
- 830
Mahalanobis reads 0.336 - worse than a coin, and nothing is broken. It measures distance from the centre, the bridge sits at the centre, so those 20 points have the LOWEST scores in the dataset: ranked below ordinary data rather than above it. It still finds 100% of the far points and 0% of the bridge. The method encodes a definition - an anomaly is far from the middle - and this data contains anomalies that are not. A detector cannot find a kind of anomaly its score function does not express.
The threshold controls recall; contamination controls precision
Price the two mistakes - say a miss costs 500 and a false alarm 20 - and the threshold stops being a matter of taste:
| policy | alerts | caught | total cost |
|---|---|---|---|
| never fire | 0 | 0 / 100 | 50,000 |
| flag the top 1% | 201 | 52 / 100 | 26,980 |
| cost-minimising | 460 | 67 / 100 | 24,360 |
The top-1% rule lands within 11% of optimal here, which is luck rather than principle: it fires less than half the alerts it should, and at a different cost ratio it would be badly wrong.
Then leave that threshold alone and let the world move:
| anomaly rate | alerts | precision | recall |
|---|---|---|---|
| half the usual | 425 | 8% | 64% |
| what it was tuned on | 460 | 15% | 67% |
| twice the usual | 523 | 25% | 65% |
Recall barely moves; precision swings threefold. Recall is the share of anomalies above the cut, so it depends on the anomaly score distribution alone, and that distribution has not changed - only the number of draws from it. The normal distribution has not changed either, which holds the false alarms near 2.3% of the normal events. So precision is a moving count of real anomalies over a nearly fixed pile of false alarms, and that ratio is what swings.
A team promised "most alerts will be real" will see that promise broken by a quiet month, with no change to the model, the pipeline or the threshold.
The figure below lets you do the thing the table only describes: hold the cut still and move the world. Drag the contamination and the recall read-out does not change by a single point, because recall reads off the anomaly score distribution alone and that distribution has not moved. Precision changes about thirteen-fold across the slider's full range, from 10 to 200 anomalies per 10,000. The capacity table underneath is the same detector reported the way it should be
- given how many items can be reviewed in a day, here is what that buys.
Interactive: hold the cut still and let the world move
The threshold never changes what share of anomalies it catches.
- Alerts
- 518
- Precision
- 13%
- Recall
- 66%
- Accuracy
- 97.56%
What each review capacity actually buys
| reviewed per day | precision | recall |
|---|---|---|
| 20 | 73% | 15% |
| 100 | 37% | 37% |
| 460 | 14% | 63% |
| 1000 | 8% | 76% |
20,000 events
At this cut the detector catches 66% of whatever anomalies exist, and that number does not depend on how many there are - drag the contamination slider and watch it refuse to move. Precision is 13% and it moves with the rate almost proportionally, because the alerts are overwhelmingly false ones and it is their ratio to the real ones that changes. So a team promised that most of what they see will be real has that promise broken by a quiet month, with no change to the model, the data or the threshold. Note the accuracy read-out while you do it: at 50.0 anomalies per ten thousand events, a detector that never fired at all would score about the same.
Specify the problem before choosing a method
Three numbers decide whether an anomaly detector is worth anything, and all three are known before any model is trained:
- the base rate - it sets what every metric can mean
- the review capacity - it fixes k, and k fixes the threshold, not the other way round
- the cost of a miss relative to a false alarm - it turns the threshold into arithmetic
Report a detector as "at forty reviews a day we catch a quarter of incidents, and three in eight of what we show you is noise". Not as "ROC 0.95".
And check which kind of anomaly your data actually contains, because every family is blind to a kind another one finds easily. Running two families with different blind spots and looking at what only one of them flags is cheap - and the disagreement is usually where the interesting anomaly is.
References & further reading
- Charu C. Aggarwal, Recommender Systems: The Textbook, Springer, 2016· Kudos AI reference library
- Kevin P. Murphy, Probabilistic Machine Learning: An Introduction, MIT Press (Adaptive Computation and Machine Learning), 2022source ↗
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.